Illumina NovaSeq: What It Is & How It Works

The Illumina NovaSeq is a high-throughput DNA sequencing instrument that reads billions of short DNA fragments in parallel, making it one of the most widely used platforms in genomics research and clinical testing. First released as the NovaSeq 6000 in 2017 and updated with the NovaSeq X series in 2023, the system relies on a method called sequencing by synthesis, where fluorescently labeled molecules are incorporated one base at a time and photographed by a built-in camera. The details of how that process works, and the quirks it introduces, matter more than most users realize.

Sequencing by Synthesis in Plain Terms

Every NovaSeq run starts with a prepared “library” of DNA fragments, each tagged with short adapter sequences on both ends. These fragments are loaded onto a glass surface called a flow cell, where they attach to complementary oligonucleotides embedded in the glass. Through a chemical amplification step, each original fragment is copied many times in a tight cluster so its signal is bright enough for the camera to detect.

Once clusters are formed, the machine floods the flow cell with four types of fluorescently tagged nucleotides, each corresponding to one of DNA’s four bases (A, C, G, T). In each cycle, one nucleotide locks into position on each growing strand, and a chemical blocker prevents the next base from attaching until the camera has taken its picture. The machine images the entire flow cell, records which color appeared at each cluster position, then strips off the blocker and the fluorescent tag, and repeats. After hundreds of cycles, the instrument has read a stretch of sequence from each cluster. This core technique is shared across Illumina’s product line; what makes the NovaSeq distinct is the scale, speed, and specific chemistry choices involved.

Patterned Flow Cells and ExAmp Chemistry

Older Illumina instruments used “random” flow cells, where DNA fragments landed wherever they happened to stick. The NovaSeq uses patterned flow cells, which have billions of tiny wells etched into the glass surface at fixed, evenly spaced positions. Each well is designed to capture one cluster. This layout lets the camera know exactly where to look, which speeds up imaging and packs far more data onto a single flow cell.

To seed those wells, the NovaSeq uses an amplification method called exclusion amplification, or ExAmp. Instead of the older bridge amplification approach, ExAmp performs the copying reaction right on the flow cell surface in a single isothermal step. The chemistry is faster and scales well, but it has a practical side effect: copies of a single original fragment can spill into neighboring wells, creating what are known as optical duplicates. These are reads that look like independent data but are really echoes of the same molecule.

In a head-to-head comparison with the Element AVITI platform, the NovaSeq X Plus showed duplicate rates of roughly 12 to 13 percent even under good loading conditions, driven primarily by optical duplicates. The AVITI, which uses a different clustering approach, kept duplicates below 5 percent and saw almost no optical duplicates at all.1PubMed Central. Whole-genome sequencing with AVITI and NovaSeq X Plus reveals comparable performance with contextual biases Bioinformatics pipelines routinely flag and remove optical duplicates, so they do not ruin the data, but they do mean a fraction of the sequencing capacity is effectively wasted on redundant reads. Labs account for this when planning how much sequencing to order.

Two-Channel Detection

One of the bigger design choices in the NovaSeq line is its two-channel optical system. Earlier Illumina instruments like the HiSeq 2500 used four separate color channels, one per base. The NovaSeq instead uses just two channels and encodes the four bases through combinations of colors: one base lights up in one channel only, another lights up in both channels, a third lights up in the other channel only, and the fourth base produces no signal at all. This setup halves the number of images the camera needs to take per cycle, which is a big part of how the instrument runs so fast.

The tradeoff is that the “dark” base, the one identified by the absence of signal, is inherently harder to call with confidence. Researchers combining data from older four-channel instruments with newer two-channel NovaSeq data have found that erroneous guanine base calls from the two-channel system can survive filtering and lead to misleading results if the data sets are naively merged.2PubMed. Sequencing platform shifts provide opportunities but pose challenges for combining genomic data sets This does not mean the NovaSeq data is unreliable on its own; it means the two detection systems have different systematic biases, and mixing them carelessly is risky.

Error Profiles and Quality Scores

Both the NovaSeq 6000 and the newer NovaSeq X report quality scores in a binned format, assigning each base call to one of just three quality tiers rather than a continuous scale. On the NovaSeq X, those bins are labeled Q12 (low confidence), Q24 (medium), and Q40 (high), with about 99 percent of bases landing in the Q40 bin. On the NovaSeq 6000, the bins are Q11, Q25, and Q37, with about 97 percent in the top tier.3bioRxiv. Comprehensive Error Profiling of NovaSeq6000, NovaSeqX, and Salus Pro Using Overlapping Paired-End Reads

In practice, the instruments are actually more accurate than those scores suggest. When researchers measured the true error rate empirically using overlapping paired-end reads, the bases labeled Q40 on the NovaSeq X turned out to perform closer to Q56, and the Q37 bases on the NovaSeq 6000 were closer to Q54. The binning system essentially rounds down, which is conservative but means the quality scores should not be taken as a precise measure of accuracy. If your pipeline uses quality scores for filtering, the coarse bins might throw away reads that are actually fine, or lump together bases of somewhat different reliability.

The two generations also make different types of errors. The NovaSeq 6000 showed a pronounced bias toward A-to-C and T-to-G substitutions, consistent with spectral cross-talk between its red and green detection channels. The NovaSeq X, which switched to blue and green channels and a new chemistry called XLEAP-SBS, shifted to a different error spectrum dominated by C-to-A and C-to-T substitutions.3bioRxiv. Comprehensive Error Profiling of NovaSeq6000, NovaSeqX, and Salus Pro Using Overlapping Paired-End Reads Earlier profiling across multiple Illumina instruments also noted that the common pattern of substituting a base of the same type as the preceding base was inconsistent on the NovaSeq 6000, with adenine over-represented in unexpected positions.4NAR Genomics and Bioinformatics. Sequencing error profiles of Illumina sequencing instruments These biases are subtle enough that they rarely matter for routine variant calling, but they can affect specialized applications like ultra-low-frequency mutation detection or methylation analysis, where even a tiny systematic error mimics a real biological signal.

Index Swapping

When labs run multiple samples on the same flow cell (a standard cost-saving practice called multiplexing), each sample’s DNA is tagged with a unique short barcode sequence, called an index. After sequencing, the software reads the index to sort each fragment back to the correct sample. On instruments that use ExAmp chemistry, including the NovaSeq, a small but measurable fraction of index sequences detach and reattach to the wrong library fragments during the clustering step. The result is that reads from one sample get misassigned to another.

In a study of hundreds of tumor and matched normal sample pairs, data generated on ExAmp instruments showed significantly higher cross-sample contamination compared to older bridge-amplification instruments, with a median contamination rate of about 0.84 percent versus 0.19 percent. Sequencing a single sample per lane eliminated the problem, but that approach is too expensive for most projects. Even rigorous bead- or gel-based purification of libraries before loading was found to be insufficient.5Scientific Reports. Sample-Index Misassignment Impacts Tumour Exome Sequencing

The most effective practical solution has been unique dual indexing, where each sample gets a distinct pair of barcodes rather than a single one. Reads that underwent index swapping will show a mismatched pair that does not correspond to any sample in the experiment and can be cleanly filtered out. Validated sets of 96 non-combinatorial dual indexes have been published and are now widely adopted across different library preparation methods.6PubMed Central. Characterization and remediation of sample index swaps by non-redundant dual indexing on massively parallel sequencing platforms Dual indexing has become standard practice for any multiplexed NovaSeq run, and newer Illumina library kits ship with it built in.

Where the NovaSeq Gets Used

The NovaSeq’s combination of high throughput and relatively low per-base cost has made it the workhorse behind several of the largest genomics projects. The expanded 1000 Genomes Project, for example, used the NovaSeq 6000 to sequence 3,202 human genomes at high coverage, discovering over 117 million small variant sites and completing 602 parent-child trios that allow researchers to study how mutations are inherited.7Cell Genomics. A high-coverage genome sequencing resource for 3,202 genomes Projects at that scale were impractical before instruments of this throughput class existed.

Single-cell RNA sequencing, where the transcriptome of each individual cell in a tissue sample is read separately, has also leaned heavily on the NovaSeq. The platform is commonly paired with droplet-based cell isolation methods to generate the billions of short reads needed to profile thousands of cells per experiment.8Quantitative Biology. Comparative analysis of NovaSeq 6000 and MGISEQ 2000 single‐cell RNA sequencing data

In clinical oncology, one of the more promising applications involves liquid biopsy, where tumor-derived DNA fragments circulating in a patient’s blood are analyzed without needing a tissue biopsy. Whole-genome sequencing of cell-free DNA on the NovaSeq has been shown to detect tumor-derived DNA at high sensitivity, distinguish between different cancer types based on epigenetic signatures, and monitor disease progression over time.9Nature Communications. Multimodal analysis of cell-free DNA whole-genome sequencing for pediatric cancers with low mutational burden A multimodal approach combining copy number, fragmentation, and methylation signals from cell-free DNA sequencing achieved around 85 percent sensitivity for detecting circulating tumor DNA across several solid cancer types.10Nature Communications. Multimodal cell-free DNA whole-genome TAPS is sensitive and reveals specific cancer signals An ultra-low-pass whole-genome sequencing approach on the NovaSeq has also been explored for risk assessment in large B-cell lymphoma, motivated by the lower cost compared to deep targeted panels.11PubMed Central. Risk Assessment With Ultra-Low-Pass Whole-Genome Sequencing of Cell-Free DNA for Large B-Cell Lymphoma

Automation and Clinical Validation

For clinical labs running the NovaSeq as part of a diagnostic workflow, reproducibility and regulatory compliance add a layer of complexity beyond what research labs typically worry about. Automated library preparation using robotic liquid-handling systems is increasingly common because it reduces hands-on time and minimizes human-introduced variation. A validation study comparing an automated NovaSeq 6000 workflow to a CE-IVD certified reference system found that the automated setup produced slightly higher duplication rates (about 15 percent versus 9 percent) but that this difference did not affect variant detection, with coverage depth and variant concordance remaining fully comparable.12PubMed Central. Validation of the NovaSeq6000 platform and automated library preparation for CE-IVD equivalence Elevated duplication from robotic workflows is a known and accepted tradeoff; labs simply plan for marginally more raw sequencing to compensate.

How It Compares to Competing Short-Read Platforms

The NovaSeq has dominated the high-throughput short-read market for years, but newer competitors have narrowed the gap. The Element AVITI, which uses a fundamentally different clustering approach that avoids the ExAmp chemistry, achieved variant-calling performance highly comparable to the NovaSeq X Plus in whole-genome sequencing, with the AVITI providing slightly better accuracy for small insertions and deletions at low coverage depths.1PubMed Central. Whole-genome sequencing with AVITI and NovaSeq X Plus reveals comparable performance with contextual biases The Ultima Genomics UG 100, which uses a different flow-cell architecture entirely, showed highly concordant results with the NovaSeq X Plus for single-cell RNA sequencing, robustly capturing all major immune cell lineages with minimal technical discrepancy between platforms.13CrossRef. Benchmarking Next-Generation Sequencing Platforms: A Comprehensive Comparison Of Single-Cell RNA-Seq from Ultima UG 100 vs. Illumina NovaSeq X Plus

The competitive picture, then, is not that the NovaSeq is technically superior to everything else. It is that the NovaSeq’s installed base, its deep integration into existing lab workflows, and the sheer volume of reference data generated on Illumina instruments create a kind of ecosystem gravity. Switching platforms means revalidating pipelines, retraining staff, and accepting that subtle systematic differences between instruments will need to be characterized and accounted for. For labs starting fresh, the newer platforms represent genuine alternatives. For labs already running NovaSeqs, the switching costs are real.

Short Reads Versus Long Reads

The NovaSeq reads DNA in fragments of roughly 150 to 250 bases, which are then computationally stitched back together. Long-read platforms from companies like Oxford Nanopore and Pacific Biosciences produce reads tens of thousands of bases long, which makes them far better at resolving repetitive regions, structural variants, and complex genomic rearrangements that short reads simply cannot span.

In a precision oncology benchmarking study, short-read Illumina data called single-nucleotide variants more accurately in homopolymers and short tandem repeats, while Oxford Nanopore long reads performed better in segmental duplications and large tandem repeats. For insertions and deletions specifically, the short-read approach was far more accurate across nearly every genomic context, with an F1 score of about 99.6 percent compared to roughly 72.5 percent for the long-read method tested.14Cell Genomics. Truth challenge V2: a precisionFDA benchmark study for variant calling in precision oncology and genomics In environmental metagenomics, short-read assemblies were less error-prone overall but struggled to assemble complicated genome regions like repeats, potentially underestimating microbial diversity in those areas. Long reads improved assembly contiguity and recovery of variable regions.15PubMed Central. Comparison of short-read and long-read metagenome assemblies in a natural soil community highlights systematic bias in recovery of high-diversity populations

Many large projects now use both: the NovaSeq provides deep, accurate coverage of the “easy” parts of the genome at low per-base cost, while long-read instruments fill in the structurally complex regions. The question of short versus long reads has largely moved past “which is better” and into “which combination is most cost-effective for this specific question.” For straightforward clinical applications like detecting known point mutations, the NovaSeq’s accuracy and throughput remain hard to beat. For de novo genome assembly or characterizing structural variants in cancer, long reads are increasingly essential.