PROTECT YOUR DNA WITH QUANTUM TECHNOLOGY
Orgo-Life the new way to the future Advertising by AdpathwayA new artificial-intelligence system developed by researchers at The University of Hong Kong could make it significantly easier to identify cancer-causing mutations hidden in the most complicated regions of the human genome. Known as ClairS, the deep-learning algorithm is designed for long-read DNA sequencing, a technology increasingly viewed as a powerful way to detect genetic changes that conventional short-read methods can overlook.
Cancer mutations are alterations in the DNA of tumour cells that are absent from healthy tissue. Finding these changes accurately is essential for understanding how cancers develop, tracking disease progression and selecting treatments tailored to individual patients. Yet the task is technically demanding. Tumour samples often contain a mixture of cancerous and normal cells, while mutations can occur in repetitive or structurally complex sections of the genome that are difficult to reconstruct from short fragments of DNA.
Most existing somatic-variant callers—the software tools used to distinguish tumour mutations from inherited genetic differences—were created primarily for short-read sequencing. Short-read platforms produce large numbers of highly accurate fragments, but each fragment covers only a small portion of the genome. Long-read sequencing, by contrast, generates much longer DNA molecules that can span repetitive sequences, structural rearrangements and other difficult regions. This broader view can reveal genomic changes that would otherwise remain hidden, although it also creates new computational challenges.
ClairS tackles these challenges with a neural-network architecture trained to interpret the complex signals produced by long-read tumour-normal sequencing. The system compares DNA data from a tumour with a matched normal sample and searches for small somatic variants, including single-nucleotide changes and short insertions or deletions. Rather than relying only on fixed rules, the model learns patterns associated with genuine tumour mutations, sequencing errors and differences caused by the proportion of cancer cells present in a sample.
One of the most innovative aspects of ClairS is the way its developers generated training data. High-quality tumour-normal datasets are scarce, expensive to produce and difficult to obtain in sufficient quantities. To overcome this limitation, the researchers mixed sequencing data from normal human samples to create synthetic tumour-normal pairs. The process allowed them to simulate a wide range of biological and technical conditions, including different tumour purities, sequencing depths and mutation burdens.
This synthetic-data strategy gives the model access to an effectively unlimited supply of realistic training examples. Tumour purity is particularly important because a mutation may appear in only a small fraction of the DNA molecules analysed. If the cancer cells represent a minor component of a biopsy, the signal from a true mutation can be overwhelmed by normal DNA. By exposing ClairS to simulated samples with varying levels of tumour purity, the researchers trained it to recognise weak but meaningful mutation signals under conditions that resemble real clinical specimens.
The team evaluated ClairS using datasets from several cancer types, including breast cancer, lung cancer, melanoma and pancreatic cancer cell lines. Across different sequencing conditions, the algorithm showed high accuracy in detecting small somatic mutations. Its performance was particularly important in regions where long reads provide an advantage, because the extended DNA fragments can preserve the genomic context needed to distinguish a true mutation from a technical artefact.
Unlike many experimental algorithms that remain confined to academic demonstrations, ClairS has already been incorporated into the official somatic-variant-calling workflow of Oxford Nanopore Technologies. The integration places the method inside a practical commercial analysis pipeline and could speed its adoption by researchers and clinical genomics laboratories. Although further validation will be needed before any tool is used routinely for patient diagnosis or treatment decisions, the development represents a significant step toward making long-read cancer analysis more accessible.
“Long-read sequencing is transforming how we study cancer genomes, especially in regions that were previously difficult to analyse,” said Professor Ruibang Luo, the study’s senior researcher and an Associate Professor at HKU’s School of Computing and Data Science. “ClairS makes it possible to train powerful AI models even when real cancer training data is limited, supporting more reliable cancer mutation discovery from long-read sequencing data.”
The work also illustrates a broader shift in biomedical AI. Many medical algorithms are limited not by a lack of computational power, but by the shortage of accurately labelled clinical data. ClairS demonstrates how carefully designed simulations can provide a practical bridge between limited real-world samples and the enormous diversity of conditions encountered in biology. By combining long-read sequencing with deep learning and scalable synthetic-data generation, the method could help researchers build more complete cancer genomes, uncover mutations missed by traditional approaches and advance the development of precision oncology. The study, published in Nature Methods, is open source, allowing the wider genomics community to inspect, reproduce and further develop the technology.
Subject of Research: Computational simulation/modeling
Article Title: ClairS: a deep-learning method for long-read tumor–normal pair somatic small variant calling
News Publication Date: 1 July 2026
Web References: https://github.com/HKU-BAL/ClairS
References: Nature Methods. DOI: 10.1038/s41592-026-03152-4
Image Credits: The University of Hong Kong
Keywords: ClairS, cancer mutations, long-read sequencing, deep learning, artificial intelligence, somatic variant calling, tumour genomics, precision medicine, bioinformatics, Oxford Nanopore Technologies
Tags: advanced variant calling methodsAI-driven cancer genomics toolscancer mutation detectioncomplex genome region analysisdeep-learning algorithms for genomicsDNA sequencing technologieslong-read DNA sequencingpersonalized cancer treatmentsomatic mutation identificationstructural genome rearrangementstumor genetic variation analysistumor heterogeneity analysis


6 hours ago
11




















English (US) ·
French (CA) ·