Next Generation Sequencing: How It Works & Why It Matters

Next generation sequencing refers to a family of technologies that read millions of DNA or RNA fragments simultaneously, producing in days or hours what once took years. The original Human Genome Project, using older Sanger sequencing, needed more than a decade to deliver a finished draft of a single human genome. Modern NGS platforms can sequence an entire human genome in roughly a day, at a fraction of the cost. That speed and scale have made NGS the backbone of genomic medicine, cancer diagnostics, prenatal screening, infectious disease surveillance, agriculture, forensics, and a growing list of fields that depend on reading genetic code quickly and affordably.

The Basic Workflow

Despite differences between platforms, virtually every NGS run follows the same general arc: extract the DNA (or RNA), prepare a library, sequence it, and analyze the data computationally. Understanding each step helps explain both the power and the limitations of the technology.

Library preparation is where raw genetic material gets converted into something the sequencer can actually read. The DNA is broken into small fragments, and short adapter sequences are attached to the ends of those fragments so the machine can grab onto them. This sounds straightforward, but the details matter enormously. A systematic comparison of library preparation kits found that overall efficiency varied by a factor of four to seven between kits, and that adaptor ligation yields alone differed by more than tenfold. When ligation efficiency is very low, the original complexity of the sample can be lost before sequencing even begins, meaning the final data may not faithfully represent what was in the original sample.1PubMed Central. Quantitation of next generation sequencing library preparation protocol efficiencies using droplet digital PCR assays – a systematic comparison of DNA library preparation kits for Illumina sequencing

Once the library is built, sequencing happens through massively parallel reading. Millions of fragment copies are anchored on a surface (a flow cell, in the most widely used Illumina platform), and the machine reads each fragment base by base, detecting signals as each nucleotide is incorporated. The result is millions of short “reads,” typically 100 to 300 bases long, each representing one small stretch of the genome.

Those raw reads then flow into a bioinformatics pipeline, where software maps them to a reference genome and calls variants, which are the places where the sample’s DNA differs from the reference. The choice of mapping and variant-calling tools matters more than many users realize. A large benchmarking study found that the GATK variant caller achieved the highest accuracy with most mapping tools, working best with alignments from BWA-MEM and Novoalign. But different combinations of mapper and caller produced meaningfully different results, and some pairings that performed well for one metric (like specificity) performed poorly on another (like overall accuracy).2bioRxiv. Comparison of read mapping and variant calling tools for the analysis of plant NGS data This means two labs analyzing the same raw sequencing data can reach slightly different conclusions depending on their software choices.

Finding the Mutations That Drive Cancer

Oncology is one of the fields most transformed by NGS. Tumors accumulate genetic mutations as they grow, and identifying those mutations can guide treatment. Traditional tissue biopsies require a surgeon to physically sample the tumor, which is invasive, sometimes risky, and only captures a single snapshot of a cancer that may be genetically diverse across different regions.

Liquid biopsy has changed this picture. Tumors shed fragments of their DNA into the bloodstream, called circulating tumor DNA (ctDNA). By drawing a simple blood sample and sequencing the cell-free DNA in it, clinicians can profile tumor mutations with minimal invasion and repeat the test as often as needed to track how the cancer evolves.3PubMed Central. Liquid Biopsy, ctDNA Diagnosis through NGS NGS-based liquid biopsy panels can assess hundreds of mutations from a small amount of DNA input, and the approach is being used across the full arc of cancer care, from initial diagnosis through treatment monitoring to detecting relapse.

The technology is not perfect, though. In a study of lung and colorectal cancer patients, NGS detected EGFR mutations in plasma with a sensitivity of about 77%, meaning it missed roughly one in four known mutations. Sensitivity was affected by whether the primary tumor was still present and by how genetically heterogeneous the cancer was.4PubMed Central. Limits and potential of targeted sequencing analysis of liquid biopsy in patients with lung and colon carcinoma For patients where tissue is unavailable, plasma-based cfDNA analysis is an established alternative for guiding treatment decisions, and NGS’s ability to test many genes at once gives it an edge over older methods that look at one mutation at a time.5PubMed Central. Next generation sequencing techniques in liquid biopsy: focus on non-small cell lung cancer patients

Diagnosing Rare and Inherited Diseases

For families navigating an undiagnosed condition in a child, NGS has been genuinely life-changing. There are thousands of known rare genetic diseases, many caused by a single mutation in a single gene. Before NGS, tracking down the right gene could take years of sequential testing. Now, exome sequencing (which reads just the protein-coding portions of the genome, roughly 1-2% of total DNA) or whole-genome sequencing (which reads everything) can survey the entire genetic landscape in one test.

A meta-analysis of pediatric rare disease studies found that genome-wide sequencing achieved a pooled diagnostic yield of about 34%, compared to roughly 18% for non-genome-wide approaches, giving it about 2.4 times the odds of reaching a diagnosis.6PubMed. A meta-analysis of diagnostic yield and clinical utility of genome and exome sequencing in pediatric rare and undiagnosed genetic diseases A large Korean study of over 1,400 families with suspected rare disorders found an even higher yield of 46%, with family-based testing (sequencing the child alongside parents) outperforming singleton testing. Neuromuscular and neurodevelopmental disorders showed the highest success rates.7npj Genomic Medicine. Clinical utility of genome sequencing in rare diseases: lessons from a single-center study of 1,452 Korean families

An important nuance is the difference between exome and whole-genome sequencing. When exome sequencing fails to find a diagnosis, whole-genome sequencing can sometimes succeed by detecting variants in non-coding regions, deep intronic sites, or complex structural rearrangements that exome panels miss. A meta-analysis estimated that genome sequencing established a diagnosis in about 7% more patients after a negative exome result. But when exome data is periodically reanalyzed using updated knowledge of gene-disease associations, reanalysis can achieve similar yields to running a whole new genome test, which highlights the value of going back to old data with fresh eyes.8PubMed. Diagnostic Yield of Genome Sequencing Versus Exome Sequencing in Pediatric Patients With Rare Phenotypes: A Systematic Review and Meta-Analysis

Prenatal Screening

One of the most widespread consumer-facing uses of NGS is noninvasive prenatal testing, or NIPT. During pregnancy, fragments of fetal DNA circulate in the mother’s blood. By sequencing that cell-free DNA, labs can screen for chromosomal conditions like Down syndrome (trisomy 21) without the miscarriage risk associated with amniocentesis or chorionic villus sampling.

A landmark trial comparing cfDNA-based screening to standard methods found dramatically lower false positive rates: 0.3% versus 3.6% for trisomy 21, and 0.2% versus 0.6% for trisomy 18. The positive predictive value for trisomy 21 was about 46% with cfDNA testing, compared to just 4% with standard screening, meaning that when cfDNA screening flags a pregnancy as high-risk, it is right far more often.9PubMed. DNA sequencing versus standard prenatal aneuploidy screening A large retrospective study of over 36,000 pregnancies confirmed these findings at scale, with NIPT results confirmed by diagnostic testing in over 99% of trisomy 21 cases, about 91% of trisomy 18 cases, and roughly 84% of trisomy 13 cases.10PubMed Central. Performance of cell-free DNA sequencing-based non-invasive prenatal testing: experience on 36,456 singleton and multiple pregnancies

NIPT is a screening test, not a diagnostic one, so a positive result still needs confirmation through invasive testing. And certain maternal conditions can affect its accuracy. Obesity, active autoimmune disease, and low-molecular-weight heparin treatment can all influence cfDNA levels in ways that change test performance.11PubMed Central. Noninvasive prenatal testing for aneuploidy using cell-free DNA – New implications for maternal health

Tracking Pathogens and Outbreaks

During the COVID-19 pandemic, many people became familiar with the idea of sequencing a virus to track its variants. That was NGS at work. But the applications in infectious disease go well beyond viral surveillance. Metagenomic NGS (mNGS) takes a clinical sample, such as cerebrospinal fluid, blood, or fluid from the lungs, and sequences everything in it without needing to know in advance what pathogen to look for. This “hypothesis-free” approach can simultaneously detect bacteria, viruses, fungi, and parasites, making it especially useful for puzzling cases where standard cultures and targeted tests come back negative.12PubMed Central. Metagenomic Next-Generation Sequencing in Infectious Diseases: Clinical Applications, Translational Challenges, and Future Directions mNGS can also detect antimicrobial resistance genes directly from the sample, giving clinicians information about which drugs are likely to work before culture results are available.

Reading Gene Activity With RNA Sequencing

DNA sequencing tells you what genetic instructions a cell carries. RNA sequencing (RNA-Seq) tells you which instructions the cell is actually using at a given moment. By capturing and sequencing the messenger RNA in a sample, researchers can measure how active each gene is, which matters for understanding disease processes, drug responses, and normal development.

RNA-Seq largely replaced older microarray technology for gene expression profiling, and for good reason. Microarrays can only detect genes for which probes have been pre-designed, and they struggle with genes expressed at low levels. RNA-Seq is better at picking up low-abundance transcripts and, critically, can distinguish between different isoforms of the same gene.13PubMed Central. Advantages of RNA-seq compared to RNA microarrays for transcriptome profiling of anterior cruciate ligament tears This isoform resolution matters clinically. For example, one comparison between the two technologies showed that a gene called RORC produces two isoforms with completely different functions: one drives inflammatory immune cell development, the other is involved in metabolism. Microarrays could not tell them apart, while RNA-Seq clearly identified which isoform was active in immune cells.14PLoS ONE. Comparison of RNA-Seq and Microarray in Transcriptome Profiling of Activated T Cells

Single-Cell and Spatial Approaches

Standard sequencing averages the signal across millions of cells in a sample, which is like measuring the average temperature of every room in a building and calling it informative. If one room is on fire and the rest are cold, the average looks normal. Single-cell RNA sequencing (scRNA-seq) solves this by isolating individual cells and sequencing each one separately, revealing the diversity hidden within what looked like a uniform tissue.

Microfluidic platforms now make it possible to process thousands of individual cells in a single experiment.15PubMed Central. Microfluidics Facilitates the Development of Single-Cell RNA Sequencing Droplet-based scRNA-seq, in which each cell is encapsulated in a tiny droplet for barcoding and lysis, has become the dominant approach and has been applied across cancer biology, developmental biology, and reproductive science.16PubMed Central. Droplet-based single-cell RNA sequencing: decoding cellular heterogeneity for breakthroughs in cancer, reproduction, and beyond Early work in this area demonstrated that enhanced measurement precision helped distinguish genuine biological variability from technical noise when studying gene expression in mouse embryonic cells.17PubMed Central. Microfluidic single-cell whole-transcriptome sequencing

Spatial transcriptomics pushes this further by measuring gene expression while preserving where in the tissue each measurement came from. Traditional sequencing destroys the tissue’s architecture during sample preparation, so you know what genes are active but not where. Spatial methods keep that map intact, making it possible to link gene expression patterns to the physical structure of tumors, organs, or developing embryos.18PubMed Central. Combining spatial transcriptomics with tissue morphology Researchers are now combining spatial gene expression data with AI-driven analysis of tissue images, enabling systems that can predict gene activity patterns from routine pathology slides.19Nature Biomedical Engineering. Integrating spatial gene expression and breast tumour morphology via deep learning

The Microbiome Connection

Your gut harbors trillions of microorganisms, and fluctuations in that microbial community correlate with conditions ranging from inflammatory bowel disease to diabetes. NGS is the primary tool for studying these communities. Two main strategies are used: 16S rRNA sequencing, which targets a single bacterial gene to identify which species are present, and shotgun metagenomics, which sequences all DNA in a sample and provides a richer picture of both species composition and metabolic capabilities.20PubMed Central. Characterization of the Gut Microbiome Using 16S or Shotgun Metagenomics

A direct comparison of the two approaches in colorectal cancer patients found that 16S captures only a portion of the community that shotgun sequencing reveals, with sparser data and lower diversity estimates. At finer taxonomic levels the two methods diverged considerably, partly because they rely on different reference databases. Still, both approaches identified microbial signatures previously linked to colorectal cancer development, and neither demonstrated clear superiority in predictive modeling for disease classification.21PubMed Central. Comparison between 16S rRNA and shotgun sequencing in colorectal cancer, advanced colorectal lesions, and healthy human gut microbiota Sequencing-based microbiome studies have also shown that many gut microbial traits are moderately to highly heritable, with heritable markers linked to conditions including type 2 diabetes, rheumatoid arthritis, and colorectal cancer.22Cell Systems. Host Genetics and Environment Shape the Human Gut Microbiome

Population-Scale Genomics

National biobank initiatives are applying NGS to hundreds of thousands of people at once, building databases that link genome sequences to health records, lifestyle data, and clinical outcomes. The UK Biobank has now whole-genome-sequenced over 490,000 participants, making it one of the largest sequencing efforts in history.23PubMed Central. Whole-genome sequencing of 490,640 UK Biobank participants Similar programs are running in the United States (All of Us), Japan, Korea, and Singapore, each generating unprecedented volumes of high-resolution genomic data integrated with phenotypic and environmental information. Key discoveries from these projects include the identification of rare genetic variants associated with complex diseases and gene expression patterns that influence disease risk across diverse populations.24PubMed Central. Lessons from national biobank projects utilizing whole-genome sequencing for population-scale genomics

Short Reads Versus Long Reads

Most widely used NGS platforms produce short reads, typically a few hundred bases at most. This works well for detecting single-nucleotide changes and small insertions or deletions. But short reads struggle with repetitive stretches of DNA and with larger structural variations like big deletions, duplications, or rearrangements, because small fragments cannot span these regions to reveal what is going on.

Long-read sequencing technologies from Pacific Biosciences and Oxford Nanopore produce reads of thousands or even tens of thousands of bases, which makes them far better at resolving structural variants in repetitive regions. About 70-80% of insertions and deletions detected by long-read methods fall in repetitive sequences, compared to only 27-58% for short reads, reflecting the difficulty short reads have in accessing those regions.25Human Genome Variation. Comparative evaluation of SNVs, indels, and structural variations detected with short- and long-read sequencing data In non-repetitive regions, short-read tools often match long-read performance for precision, so the advantage is specifically about what short reads cannot reach.

This has real clinical consequences. In cancer genomes, long-read sequencing detected larger numbers of structural variants than short-read studies had found, and the mechanisms generating those somatic variants turned out to differ from those producing normal germline deletions.26PubMed Central. Whole-genome sequencing with long reads reveals complex structure and origin of structural variation in human genetic variations and somatic mutations in cancer For rare genetic disorders, long-read platforms can now accurately find structural variants in previously unreachable areas of the genome, including repetitive sequences and segmental duplications, which are known to harbor disease-causing mutations that standard short-read testing misses entirely.27PubMed Central. Long-Read Sequencing and Structural Variant Detection: Unlocking the Hidden Genome in Rare Genetic Disorders

The Problem of Variants of Uncertain Significance

One of the most frustrating consequences of sequencing so much DNA is that you inevitably find variants that nobody knows how to interpret yet. These are called variants of uncertain significance, or VUS, and they are a growing headache in clinical genomics. A VUS is a genetic change that has been identified in a patient but for which the evidence is not yet strong enough to classify it as either harmful or benign.

Only a minority of VUS are later reclassified as pathogenic, and that resolution is rarely timely. In the meantime, the uncertainty complicates clinical decisions and can lead to real harms, including unnecessary treatment, additional invasive testing, and psychological distress for patients and families.28PubMed Central. The Challenge of Genetic Variants of Uncertain Clinical Significance: A Narrative Review As sequencing becomes more common and broader panels are used, the number of VUS findings grows, so this is a problem that is getting larger, not smaller, with the spread of the technology. It is one of the clearest examples of how more data does not automatically mean more clarity.

Beyond Human Medicine

NGS has reshaped plant breeding by enabling rapid identification of genetic markers linked to desirable traits like yield, stress resistance, and nutritional quality. Whole-genome resequencing and high-resolution mapping of complex traits have accelerated the development of improved crop varieties, compressing timelines that once stretched over many breeding cycles.29PubMed Central. Next generation sequencing technologies for next generation plant breeding

In forensics, NGS has expanded what investigators can extract from degraded DNA samples. Traditional forensic profiling relies on short tandem repeat (STR) markers, which require relatively intact DNA. When samples are badly degraded, as they often are at crime scenes or in disaster victim identification, STR analysis may fail. NGS-based approaches can instead profile single-nucleotide polymorphisms at high resolution, improving outcomes in difficult cases.30PubMed Central. Analysis of Human Degraded DNA in Forensic Genetics Ancient DNA researchers face similar challenges with degraded material and have adopted many of the same NGS workflows, using parallel sequencing of millions of tiny DNA molecules to reconstruct genomes from archaeological remains that would have been impossible to study with older methods.31iScience. Bridging computational workflows for degraded DNA: A comparative review of forensic genetics and ancient DNA

Epigenomics and Chemical Marks on DNA

Your DNA sequence is not the whole story. Chemical modifications to DNA, particularly the addition of methyl groups to cytosine bases, influence which genes are turned on or off without changing the underlying code. These epigenetic marks play roles in development, aging, and diseases including cancer. NGS-based bisulfite sequencing is considered the gold-standard method for detecting DNA methylation at single-base resolution. The chemistry is elegant in its simplicity: treating DNA with bisulfite converts unmodified cytosines to uracil (read as thymine after amplification), while methylated cytosines resist the conversion and remain as cytosines. By comparing the treated sequence to the reference, every methylated position becomes visible.32PubMed Central. DNA methylation detection: Bisulfite genomic sequencing analysis This kind of analysis would be essentially impossible without the throughput of NGS, because you need to read the same regions many times to accurately quantify methylation levels at each position.