Sequencing by synthesis is a method of reading DNA one base at a time by watching a polymerase enzyme build a new strand. Instead of reading a finished molecule, it records each nucleotide (A, C, G, or T) as it gets added to a growing copy of the template DNA. The approach powers the vast majority of high-throughput sequencing done today, and its development at the University of Cambridge in the late 1990s eventually led to commercial instruments that have driven down the cost of reading a human genome by several orders of magnitude. The chemistry is clever, the engineering around it is enormous, and the error patterns it produces shape how researchers design and interpret experiments across nearly every branch of biology and medicine.
The Core Idea Behind the Chemistry
DNA polymerase is the enzyme cells use to copy their own genomes. It grabs free nucleotides floating in solution and snaps them into place on a growing strand, matching each one to the template: A pairs with T, C pairs with G. In normal replication this happens at blazing speed, thousands of bases per second, which is great for a cell but useless for anyone trying to read the sequence. The trick behind sequencing by synthesis is forcing the polymerase to pause after every single base so a detector can record which nucleotide just got added.
The pause is accomplished with modified nucleotides called reversible terminators. Each of the four nucleotide types carries two modifications. First, a small chemical group caps the 3ʹ-OH position, the spot on the sugar where the next nucleotide would normally attach. With that cap in place, the polymerase has nowhere to connect another base, so the reaction stalls after one incorporation. Second, a fluorescent dye is attached to the nucleotide through a cleavable linker, giving each of the four bases a distinct color.
After the polymerase adds one modified nucleotide and stops, a detector reads the fluorescent signal to identify which base was incorporated. Then a chemical treatment removes both the fluorescent dye and the 3ʹ cap simultaneously. Researchers showed early on that an allyl-based capping group, paired with a matching allyl linker for the fluorophore, could be stripped in about 30 seconds using a palladium-catalyzed reaction in water-based buffer. That single cleanup step restores the growing strand to a state where the polymerase can add the next base, and the cycle repeats.
1PubMed Central. Four-color DNA sequencing by synthesis using cleavable fluorescent nucleotide reversible terminatorsThe feasibility of this cycle was established in proof-of-concept experiments using a modified uracil nucleotide with the same allyl cap at the 3ʹ position and a photocleavable fluorophore. After the nucleotide was incorporated and the fluorescence recorded, light exposure broke the dye away, and a palladium reaction removed the cap, allowing the polymerase to continue.
2PubMed Central. Design and synthesis of a 3′-O-allyl photocleavable fluorescent nucleotide as a reversible terminator for DNA sequencing by synthesisThe Engineered Polymerase
Normal DNA polymerases did not evolve to handle bulky modified nucleotides, so the enzymes used in commercial SBS instruments are heavily engineered. One widely used variant, based on a polymerase originally isolated from a heat-loving archaeal organism, carries three specific mutations that allow it to accept a broad range of modified nucleotide substrates, including fluorescently labeled reversible terminators.
3PubMed Central. Therminator DNA Polymerase: Modified Nucleotides and Unnatural SubstratesGetting the enzyme right matters enormously for data quality. If the polymerase occasionally skips a modified nucleotide or incorporates the wrong one, those errors propagate through the rest of the read. The interplay between enzyme fidelity, the exact structure of the reversible terminator, and the chemistry used to remove the cap after each cycle are all tightly optimized in modern instruments. A 2009 review catalogued the major technical hurdles of SBS as including enzyme-substrate optimization, fluorescent label design, surface chemistry, optics, and the tradeoff between read length and accuracy.
4Nature. The challenges of sequencing by synthesisFrom Single Molecules to Clusters on a Flow Cell
A single fluorescent signal from one molecule is far too faint for most cameras to detect reliably. To solve this, SBS platforms amplify each template molecule into a dense cluster of identical copies, all anchored to a glass surface called a flow cell. In the original approach, short adapter sequences on each end of a DNA fragment bind to complementary oligonucleotides fixed to the flow cell surface. The fragment bends over, its other end hybridizes to a nearby surface oligo, and the polymerase copies it. This bridge amplification cycle repeats until each original molecule has produced a tight cluster of roughly a thousand clones, all in the same small spot.
A newer method, called exclusion amplification, was introduced in 2015 on higher-capacity instruments. Instead of randomly distributed clusters on a flat surface, the flow cell contains millions of tiny nanowells at fixed, known positions. Each nanowell ideally captures one template molecule, which is then amplified in place. The patterned layout increases the density of usable clusters, reducing the per-run cost, and improves data quality for longer reads.
5bioRxiv. Index switching causes “spreading-of-signal” among multiplexed samples in Illumina HiSeq 4000 DNA sequencingRegardless of which amplification method is used, the result is the same: millions of clusters, each glowing with its own color at each cycle. A camera images the entire flow cell after every nucleotide incorporation, and software maps each cluster’s position and color to determine the base sequence.
Quality Scores and What They Mean
Every base call produced by an SBS instrument comes with a quality score, expressed on the Phred scale. A Phred score of 30 means the instrument estimates a 0.1% chance the call is wrong, which translates to 99.9% accuracy. A score of 20 means 99% accuracy, and a score of 40 means 99.99%.
6PubMed Central. Estimating Phred scores of Illumina base calls by logistic regression and sparse modelingThese scores are not constants across a read. The first few bases tend to be noisier because the chemistry has not yet settled into a steady state. Quality also degrades toward the end of longer reads as the clusters gradually lose synchrony, a phenomenon called phasing. If a few copies in a cluster fall behind or jump ahead by a base, the fluorescent signal from that cluster becomes a mixture of colors rather than a clean single hue, and the software’s confidence in its call drops. This phasing effect is one of the main reasons SBS reads have traditionally topped out at a few hundred bases, far shorter than older Sanger sequencing reads.
Where the Errors Tend to Show Up
SBS errors are not random. They follow patterns tied to the local sequence context and the GC content of the DNA being read. A study examining Illumina sequencing data across multiple bacterial genomes found that certain base-substitution errors were far more common than others. Conversions from A or T to G or C happened more often than the reverse, regardless of species. Among those, T-to-G and A-to-C miscalls were especially frequent, likely reflecting optical crosstalk between fluorescent channels.
7PubMed Central. Sequence-specific error profile of Illumina sequencersErrors also cluster in space. An evaluation of HiSeq data found that just 3% of error-prone positions accounted for about a quarter of all substitution errors, meaning a small number of genomic spots were problematic over and over again rather than errors being sprinkled evenly. Insertions and deletions were rare on average but spiked to around 2% in homopolymer stretches, where the same base repeats many times in a row.
8PubMed Central. Evaluation of genomic high-throughput sequencing data generated on Illumina HiSeq and genome analyzer systemsGC bias is a broader issue that affects not just error rates but coverage itself. Workflows using MiSeq and NextSeq instruments showed major coverage distortions outside the 45–65% GC range. Genomic windows sitting at around 30% GC had more than ten times less coverage than windows near 50% GC.
9GigaScience. GC bias affects genomic and metagenomic reconstructions, underrepresenting GC-poor organismsFor researchers working with organisms or genomic regions that sit at extreme GC levels, this bias can make certain stretches effectively invisible in the data. It is one of the most important practical limitations to keep in mind when designing an experiment around SBS platforms.
Index Hopping and Multiplexing Artifacts
Modern SBS runs routinely sequence many different samples at once on the same flow cell, a practice called multiplexing. Each sample’s DNA fragments are tagged with a unique short barcode sequence (an index) during library preparation, and after sequencing, software sorts reads back to the correct sample based on that barcode. The system works well, but it is not perfect.
Index hopping occurs when a barcode from one sample gets attached to a DNA fragment from a different sample during the amplification or sequencing process. Early reports flagged alarmingly high misassignment rates, up to 10%, on instruments using the newer exclusion amplification chemistry. More careful analysis, controlling for other sources of noise, pinned the actual average rate at about 0.47% of reads.
10PubMed. Index hopping on the Illumina HiseqX platform and its consequences for ancient DNA studiesHalf a percent may sound trivial, but the effect is not evenly distributed. Samples that contribute few reads to the total pool receive a disproportionately high fraction of misassigned reads, because the absolute number of hopped reads is roughly proportional to the dominant sample’s contribution. For applications like ancient DNA analysis, where authentic molecules are scarce and contaminants abundant, even a small rate of index hopping can seriously distort results.
In single-cell RNA sequencing, where thousands of cells are barcoded and pooled, index hopping creates phantom gene expression. A computational correction method estimates that subtracting about 0.65% of a library’s average expression level from each cell’s count for each gene can compensate for the artifact.
11PubMed Central. Estimating and correcting index hopping misassignments in single-cell RNA-seq dataThe practical workaround most labs use is dual indexing: tagging each sample with two unique barcodes instead of one, so that a hopped read is overwhelmingly likely to carry a mismatched pair that the software can flag and discard.
How SBS Compares to Other Sequencing Approaches
SBS is not the only way to read DNA. Two major alternatives take fundamentally different approaches. Ion semiconductor sequencing, developed by Ion Torrent, also builds new strands one base at a time but skips the fluorescent labels entirely. Instead, it detects the hydrogen ion released each time a nucleotide is incorporated, using a chip packed with over 165 million tiny sensor wells. The instrument floods the chip with one nucleotide type at a time in rotation; if that base is the correct next base for a given template, a hydrogen ion is released and the pH change registers on the sensor.
12PubMed Central. Semiconductor Sequencing of Human Exomes on the Ion Proton SystemThis electronic detection is fast and eliminates the need for expensive optics, but it struggles with homopolymers. When several identical bases appear in a row, the sensor has to distinguish whether two, three, or four hydrogen ions were released in one burst, and that analog measurement is inherently less precise than the digital one-at-a-time approach of fluorescent reversible terminators.
Nanopore sequencing, developed by Oxford Nanopore Technologies, takes a completely different path. It threads a single strand of DNA through a protein pore embedded in an electrical membrane and reads the sequence from fluctuations in the current as different bases pass through. Nanopore reads can be extremely long, sometimes exceeding tens of thousands of bases, which is useful for resolving repetitive regions and structural variants that short SBS reads cannot span. A head-to-head comparison of SBS and nanopore sequencing for cell-free bacterial DNA found that SBS generated substantially more total reads and higher per-base quality, while nanopore produced longer reads. The two platforms also sometimes gave different pictures of microbial composition and diversity in the same sample.
13Scientific Reports. Comparison of Oxford Nanopore and Sequencing-by-Synthesis Technologies for Sequencing cell-free bacterial DNAIn practice, the choice between platforms often depends on what the experiment needs. SBS dominates when you want very high accuracy across billions of short reads, which suits applications like variant calling, gene expression quantification, and large-scale population studies. Long-read technologies shine for genome assembly, resolving complex structural rearrangements, and full-length transcript analysis.
SBS in Single-Cell Sequencing
One of the most transformative applications of SBS has been single-cell sequencing, where the goal is to read the transcriptome or genome of individual cells rather than a bulk mixture of millions. Platforms like 10x Genomics capture single cells in droplets, tag each cell’s RNA with a unique barcode, reverse-transcribe the RNA into complementary DNA, and then sequence everything on an SBS instrument. The result is a gene expression profile for each cell, revealing the diversity of cell types within a tissue.
A recent comparison sequenced the same single-cell libraries on both an Illumina short-read SBS platform and a PacBio long-read platform, matching individual molecules by their cell barcodes and unique molecular identifiers. The study found that both methods recovered highly comparable results in terms of cells detected and transcripts captured.
14PubMed. Comparison of single-cell long-read and short-read transcriptome sequencing via cDNA molecule matching: quality evaluation of the MAS-ISO-seq approachShort-read SBS remains the workhorse for single-cell studies because its per-read cost is low and the throughput is enormous, but long-read follow-up is increasingly used to resolve transcript isoforms that short reads cannot distinguish. Some labs now run both in tandem on the same library.
Synthetic Long Reads
A creative workaround for the short-read limitation of SBS is synthetic long-read sequencing. Rather than physically reading long stretches in one go, this approach pools small numbers of full-length complementary DNA molecules (typically a thousand or fewer per pool, often just one molecule per gene), amplifies and fragments them, and sequences the fragments on a standard short-read SBS instrument. Because each pool is so sparse, computational algorithms can reconstruct which short fragments came from the same original molecule and stitch them together into a virtual long read.
15Nature Biotechnology. Comprehensive transcriptome analysis using synthetic long-read sequencing reveals molecular co-association of distant splicing eventsThis technique lets researchers study things like distant splicing events on the same transcript, which require knowing the full-length sequence, while still using the high accuracy and low cost of SBS. It is a good example of how bioinformatics and clever library preparation can extend a platform’s capabilities beyond what the raw read length would suggest.
Clinical Uses Built on SBS
The high throughput and relatively low cost of SBS have made it the backbone of several routine clinical tests. Non-invasive prenatal testing is one of the most widely deployed. A pregnant person’s blood contains small fragments of fetal DNA circulating alongside their own. By performing ultra-low-pass SBS, reading the genome at just 0.1 to 0.3 times average coverage, laboratories can detect extra copies of chromosomes like 21, 18, or 13, screening for conditions like Down syndrome without an invasive procedure. Research has shown that even at these very shallow depths, allele frequencies can be estimated accurately enough to achieve genotype imputation accuracy above 0.84.
16PubMed. Utilizing non-invasive prenatal test sequencing data for human genetic investigationCancer genomics is another major area. Tumor sequencing panels use targeted SBS to read hundreds of cancer-associated genes at very high depth, sometimes thousands of reads per position, to find mutations present in a small fraction of tumor cells. Liquid biopsy approaches extend this further by sequencing cell-free tumor DNA from a blood draw, using the same high-depth SBS strategy to detect minimal residual disease or monitor treatment response without a tissue biopsy.
Pharmacogenomics, infectious disease surveillance, rare disease diagnosis through exome or whole-genome sequencing, and carrier screening for reproductive planning all lean heavily on SBS as their sequencing engine. The technology is so embedded in clinical genomics that for many practitioners it has become essentially synonymous with “sequencing.”
How Solexa Became Illumina
The intellectual foundations of modern SBS trace back to exploratory research at the University of Cambridge in the mid-to-late 1990s. The researchers who developed the core concepts founded a company called Solexa in 1998 to commercialize the technology. Solexa built the first commercial SBS sequencing system, and Illumina acquired the company in early 2007.
17Clinical Chemistry. Solexa Sequencing: Decoding Genomes on a Population ScaleSince the acquisition, Illumina has iterated on the platform aggressively, releasing a succession of instruments with higher cluster densities, faster run times, longer reads, and improved chemistry. The company held a near-monopoly on short-read sequencing for over a decade. More recently, competitors including MGI Tech (using a different implementation of SBS called DNA nanoball sequencing) and Ultima Genomics have entered the market, applying pressure on per-base costs. The underlying principle, reading bases one at a time as a polymerase adds them to a growing strand, remains the same across all of these platforms, even as the specific implementations of cluster generation, terminator chemistry, and signal detection diverge.
Practical Limits That Shape Experimental Design
If you are planning an experiment that will use SBS, a few constraints are worth building into your design from the start. Read length is the most obvious: most Illumina instruments top out at around 300 bases per read in paired-end mode, meaning each fragment is read from both ends. That is plenty for variant calling, gene expression, and many other applications, but it cannot span long repetitive elements or resolve complex structural rearrangements in a single read.
GC bias, as described earlier, means that organisms or regions with extreme base composition will be underrepresented. If your target genome has GC content well below 40% or above 65%, you should expect uneven coverage and may need to sequence more deeply or use PCR-free library preparation to reduce amplification bias.
Index hopping matters whenever you multiplex samples, which is nearly always. Using unique dual indexes rather than combinatorial single indexes is now considered best practice. And for applications where even a small percentage of cross-contamination is unacceptable, like ancient DNA or low-input clinical samples, additional computational filtering or running samples on separate lanes may be necessary.
Finally, the error profile’s dependence on sequence context means that systematic errors can masquerade as real variants if you are not careful. Most bioinformatics pipelines account for this by requiring a variant to appear on reads from both strands and at a minimum frequency threshold, but awareness of the underlying pattern helps when troubleshooting unexpected calls.