Sanger sequencing and Illumina sequencing read DNA using fundamentally different strategies, and those differences ripple outward into cost, speed, accuracy, and the kinds of questions each technology can answer. Sanger sequencing reads a single stretch of DNA at a time and produces long, highly accurate reads, while Illumina platforms sequence millions of short fragments simultaneously, generating vast amounts of data in a single run. The choice between them is rarely either/or in modern labs, though, and understanding where each excels reveals why both remain in active use decades after Sanger’s method first appeared.
How Sanger Sequencing Works
Sanger sequencing, developed in the late 1970s by Frederick Sanger and colleagues, relies on a clever chemical trick. A short primer is attached to a single-stranded DNA template, and a polymerase enzyme begins building a complementary strand by adding nucleotides one at a time. Mixed into the reaction are modified nucleotides called dideoxynucleotides, which look enough like normal nucleotides that the polymerase incorporates them, but they lack the chemical group needed to attach the next nucleotide. Whenever a dideoxynucleotide gets added, the growing chain stops right there.1PubMed. DNA sequencing by the dideoxy method By running millions of these reactions at once, each terminating at a different position, the result is a collection of fragments of every possible length. Sorting those fragments by size and reading which terminating nucleotide sits at the end of each one spells out the DNA sequence. Modern Sanger instruments use fluorescently labeled terminators and capillary electrophoresis to automate this, producing reads that typically run 700 to 900 bases long.
How Illumina Sequencing Works
Illumina’s approach belongs to a family of methods called sequencing by synthesis, but it scales up massively by running the chemistry on millions of tiny clusters of identical DNA fragments attached to a glass surface called a flow cell. Each cluster originated from a single fragment of the sample DNA, amplified in place until there are enough copies to generate a detectable signal. During sequencing, fluorescently labeled nucleotides are washed across the flow cell one cycle at a time. Each nucleotide carries a reversible chemical block that prevents more than one base from being added per cycle, so the instrument can photograph which color lights up at each cluster position before removing the block and starting the next cycle.2PubMed Central. The history and advances of reversible terminators used in new generations of sequencing technology This cycle-by-cycle imaging across millions of clusters is what gives Illumina its extraordinary throughput. Individual reads are short, usually between 75 and 300 bases depending on the instrument and run configuration, but the sheer number of them compensates by covering the same genomic region many times over.
Throughput and Speed
The scale difference between the two platforms is enormous. A single Sanger run on a 96-capillary instrument produces roughly 96 reads, each under a kilobase long. That is useful for confirming a handful of known targets but impractical for surveying an entire genome. Early sequencing projects, including the Human Genome Project, relied on Sanger chemistry but required years of effort and enormous resources. Today’s Illumina instruments can sequence the entire DNA of even the most complex organisms in about a day.3PubMed Central. DNA Sequencing Methods: From Past to Present A high-end Illumina machine can produce hundreds of billions of bases in a single run. That makes whole-genome sequencing, whole-exome sequencing, and large-scale population studies feasible in ways Sanger never could be.
Cost Per Base vs. Cost Per Experiment
Illumina’s cost advantage in raw sequencing data is staggering. Comparative analyses have placed the cost of sequencing on Illumina platforms at roughly ten cents or less per megabase of data, while older platforms and Sanger-based approaches cost orders of magnitude more per base.4PubMed. Field guide to next-generation DNA sequencers For whole-genome sequencing on clinical-grade Illumina instruments, a detailed cost analysis in a German hospital setting estimated the all-in price of a single genome at roughly €3,860 on an older HiSeq model and about €1,410 on the newer HiSeq X Ten, a reduction of around 63 percent driven largely by higher throughput and lower reagent costs per sample.5PubMed. Cost analysis of whole genome sequencing in German clinical practice Those numbers have continued falling in the years since.
But cost per base does not tell the whole story. If you need to check one gene variant in one patient sample, a Sanger run is fast, cheap, and uncomplicated. The library preparation, bioinformatics pipeline, and data storage overhead that come with an Illumina run can be overkill for a single-target question. Sanger sequencing remains the more economical choice when the question is narrow and the number of targets is small. The crossover point, where Illumina becomes cheaper overall, arrives once you need to interrogate dozens or more regions from the same sample.
Accuracy and Error Profiles
Both technologies are highly accurate, but they make different kinds of mistakes. In a head-to-head comparison that tested Sanger sequencing against two different Illumina exome-sequencing pipelines on the same samples, Sanger achieved about 99 percent sensitivity for detecting true variants, while the Illumina experiments ranged from 97 to 100 percent depending on the capture kit used. False-positive rates were vanishingly small for both, on the order of a few per million bases, with Sanger at about 3.7 per million and the two Illumina protocols at roughly 2.5 and 5.2 per million.6PubMed. Performance comparison: exome sequencing as a single test replacing Sanger sequencing The overall concordance was high, meaning for the vast majority of positions the two technologies agreed completely.
Where the differences matter most is not in average accuracy but in the types of errors each platform tends to produce. Sanger reads are long enough that the context around a variant is visible, and the electropherogram (the raw trace of fluorescent peaks) gives a human reviewer a visual sense of data quality at every position. Illumina’s short reads sometimes struggle with repetitive regions of the genome where fragments can be mapped to the wrong location, or with certain insertion and deletion events that are hard to call from reads only a couple hundred bases long.
Where Short Reads Fall Short
Illumina’s reliance on short reads creates blind spots that matter in clinical genetics. Certain structural features of the genome, including repetitive elements and mobile DNA insertions, can be invisible to short-read sequencing. A recent case in ophthalmic genetics highlighted this limitation: short-read methods, including both a targeted gene panel and whole-exome sequencing, failed to detect a specific Alu insertion in the RP1 gene that was associated with inherited retinal disease.7PubMed. Limitations of short-read NGS in detecting RP1 Alu insertions: a case emphasizing Sanger confirmation Sanger sequencing, with its longer reads, was able to confirm the insertion. Cases like this are uncommon but clinically significant, because a missed variant can mean a missed diagnosis.
More broadly, any genomic region where the sequence is highly repetitive or structurally complex poses a challenge for short-read alignment algorithms. The reads are so short that they can map equally well to multiple locations, creating ambiguity that the software must resolve, sometimes incorrectly. This is one reason why next-generation sequencing methods in general, despite their power, have been described as having their “most notable drawback” in short read length.8Trends in Genetics. Sanger vs. Illumina: Key Differences in DNA Sequencing
Why Sanger Validation of Illumina Results Persists
Given those blind spots, clinical laboratories routinely use Sanger sequencing to double-check variants identified by Illumina before reporting them to patients. A large validation study examined 945 variants identified by next-generation sequencing and then independently confirmed by Sanger. Only three showed any discrepancy between the two methods, and in all three cases a deeper look at the data confirmed that the original Illumina call was actually correct.9PubMed Central. Sanger Validation of High-Throughput Sequencing in Genetic Diagnosis: Still the Best Practice? That finding has led some researchers to question whether Sanger validation is still strictly necessary in every case, especially as Illumina quality metrics and bioinformatics filtering continue to improve.
Still, the practice persists in most clinical genetics labs, partly because of regulatory expectations and partly because the stakes of a false result in a diagnostic setting are so high. When a sequencing result might guide surgery, drug selection, or reproductive counseling, the extra cost of a confirmatory Sanger run is trivial compared to the consequences of an error. The question is less about whether Sanger is more accurate on average and more about whether the specific variant in question falls in a region where Illumina’s short reads might struggle.
How the Two Platforms Work Together in Practice
Modern genomics projects rarely use one sequencing technology in isolation. A study investigating a novel genetic disease illustrates the typical workflow: researchers performed whole-genome sequencing on one platform, whole-exome sequencing on an Illumina Genome Analyzer, SNP genotyping on an Illumina array, RNA sequencing on an Illumina HiSeq, and Sanger sequencing to confirm the candidate mutations that emerged.10PubMed Central. Whole-genome DNA/RNA sequencing identifies truncating mutations in RBCK1 in a novel Mendelian disease with neuromuscular and cardiac involvement Each technology contributed something the others could not. Illumina provided the broad sweep across the genome, identifying candidate genes and expression patterns. Sanger provided targeted, high-confidence confirmation of the specific mutations that mattered most.
This division of labor extends across clinical and research settings. Illumina handles discovery and screening; Sanger handles confirmation and targeted follow-up. A genetic testing lab might use Illumina to scan hundreds of genes at once and then use Sanger to verify the handful of variants that look clinically relevant. A microbiology lab might use Illumina for metagenomic surveys of environmental samples and Sanger to identify a specific pathogen isolate.
Data Formats and Bioinformatics Overhead
The data that come off each platform look quite different and demand different levels of computational effort. Sanger produces a chromatogram trace file for each read, which a trained technician can often interpret visually. The bioinformatics pipeline is minimal: align the read to a reference, call the bases, and compare to the expected sequence. Illumina data arrives as enormous FASTQ files, a format that pairs each base with a quality score, and the Sanger and Illumina variants of this format are not even directly interchangeable without conversion.11PubMed Central. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants Processing Illumina data requires dedicated computing infrastructure, specialized alignment and variant-calling software, and trained bioinformaticians. For a small lab running a handful of Sanger reactions, this overhead does not exist. For a large center running dozens of Illumina genomes per week, the computational cost is a real line item that rivals the sequencing reagent cost itself.
Storage is another consideration. A single Illumina whole-genome sequencing run at standard depth generates roughly 100 to 200 gigabytes of raw data, and retaining that data for clinical records or future reanalysis requires robust storage infrastructure. A Sanger trace file for one read is measured in kilobytes. Labs that transition from Sanger-only workflows to Illumina-based testing often underestimate the IT investment required to handle the data flood.
Where Third-Generation Sequencing Fits In
Both Sanger and Illumina represent earlier waves of sequencing technology, and a third generation has arrived with platforms from Pacific Biosciences (PacBio) and Oxford Nanopore Technologies. These instruments produce much longer reads, sometimes tens of thousands of bases or more, and they sequence single molecules of DNA without the amplification step that Illumina requires. Long-read sequencing has shown particular promise for assembling genomes from scratch, especially in regions where short reads get confused. A comparison across platforms found that PacBio’s long reads had no coverage bias in AT-rich regions, a known weakness of some short-read methods, and produced assemblies of higher structural accuracy.12PubMed Central. A precise chloroplast genome of Nelumbo nucifera (Nelumbonaceae) evaluated with Sanger, Illumina MiSeq, and PacBio RS II sequencing platforms: insight into the plastid evolution of basal eudicots
Third-generation platforms have not replaced Illumina or Sanger, though. Their per-base error rates, while improving rapidly, have historically been higher than Illumina’s, and their cost per base remains higher for large-scale projects. What they bring is read length and the ability to span structural variants, repetitive regions, and epigenetic modifications that neither Sanger nor Illumina can fully resolve. In many cutting-edge projects, all three generations of sequencing are used in combination: Illumina for deep, accurate coverage; long reads for structural scaffolding; and Sanger for pinpoint confirmation of specific variants.
Choosing Between Them
The practical decision of whether to use Sanger or Illumina comes down to the scope of the question. If you need to check a known mutation in a family member after a genetic diagnosis has already been made, Sanger is fast, cheap, and perfectly suited to the task. If you need to screen all protein-coding genes in a patient with an undiagnosed condition, exome sequencing on an Illumina platform is the clear choice. If you need to sequence a microbial genome from scratch, Illumina gives you the depth and Sanger or long-read sequencing can fill in the gaps.
Several rules of thumb help guide the decision:
- Number of targets: One to a few genes favors Sanger. Dozens or more favors Illumina.
- Prior knowledge: When you know exactly what you are looking for, Sanger is efficient. When the search is open-ended, Illumina’s breadth wins.
- Turnaround time: A single Sanger reaction can be completed in hours. An Illumina run typically takes one to three days of instrument time, plus library preparation and data analysis.
- Structural complexity: If the region of interest is repetitive or prone to insertions and deletions, short Illumina reads may miss what longer Sanger reads or third-generation reads can catch.
Common Misconceptions
One widespread misunderstanding is that Illumina has made Sanger obsolete. In research settings, Sanger volumes have certainly declined as whole-genome and whole-exome sequencing have become routine. But in clinical diagnostics, Sanger remains embedded in workflows as a confirmation tool, and for single-target assays it is still the most practical option. Labs worldwide run millions of Sanger reactions every year.
Another misconception is that Illumina sequencing is always more accurate because it generates more data. Depth of coverage does improve confidence, but the nature of the errors matters. Illumina’s systematic biases in certain sequence contexts, like homopolymer runs (stretches of the same base repeated) or GC-rich regions, are not fixed by adding more reads. They are inherent to the chemistry. Sanger reads through many of these contexts without difficulty, which is part of why it remains the confirmation standard.
A third common confusion involves equating “next-generation sequencing” with Illumina specifically. Illumina dominates the short-read sequencing market, but NGS as a category also includes other platforms with different chemistries. When a clinical report says a variant was detected by NGS, the specific platform and its known limitations matter for interpreting the result. Not all NGS is created equal, and the validation data generated on one Illumina instrument model do not automatically transfer to another.