Whole genome sequencing turns an organism’s complete DNA into readable data through a pipeline that stretches from wet-lab sample handling to computational analysis and clinical or research interpretation. The workflow is often described in three broad phases: library preparation and sequencing, bioinformatic processing, and variant interpretation. Each phase has its own decision points, and the choices made early on ripple through every downstream step. Getting the full picture of this pipeline matters whether you are setting up a new lab, evaluating a clinical report, or simply trying to understand what happens between a blood draw and a list of genetic variants.
Extracting and Fragmenting the DNA
Every sequencing run begins with high-quality DNA. For short-read platforms, most commercial extraction kits work well enough, but long-read sequencing is far more demanding. Long-read experiments require high concentrations of highly purified, high-molecular-weight DNA, which limits the usefulness of kits originally designed for short-read work.1PubMed Central. High molecular weight DNA extraction strategies for long-read sequencing of complex metagenomes Shearing forces during extraction, freeze-thaw cycles, and even aggressive vortexing can break long molecules into fragments too small for platforms that thrive on multi-kilobase reads. Specialized protocols using gentle lysis and gravity-based purification help preserve those long stretches of intact DNA.
Once extracted, the DNA usually needs to be broken into smaller pieces for sequencing. Two main approaches dominate: mechanical fragmentation (typically ultrasonication) and enzymatic fragmentation (using transposases or endonucleases). Research comparing the two shows that mechanical fragmentation yields more uniform coverage across different sample types and across the GC spectrum, while enzymatic workflows tend to produce more pronounced coverage imbalances in high-GC regions.2PubMed Central. Optimization of DNA Fragmentation Techniques to Maximize Coverage Uniformity of Clinically Relevant Genes Using Whole Genome Sequencing That unevenness can reduce the sensitivity of variant detection, especially at lower sequencing depths. Newer enzymatic library prep chemistries have narrowed the gap, though. Kits like Nextera Flex and QIAseq FX were found to be less sensitive to variable GC content than older enzymatic methods, producing a nearly uniform distribution of read depth.3Frontiers in Public Health. Evaluation of Rapid Library Preparation Protocols for Whole Genome Sequencing Based Outbreak Investigation
Library Preparation Choices
After fragmentation, adapter sequences are ligated to the DNA fragments so the sequencing instrument can grab onto them. This step, together with optional size selection and amplification, constitutes “library preparation.” One important fork in the road is whether or not to include PCR amplification. PCR-free protocols avoid the biases that amplification introduces, particularly in regions with extreme base composition. For genomes with very high AT content, PCR-free methods dramatically improve coverage uniformity compared to amplified libraries.4PubMed Central. Improvement of PCR-free NGS Library Preparation to Obtain Uniform Read Coverage of Genome with Extremely High AT Content The trade-off is that PCR-free workflows demand more starting DNA, which can be a problem when sample material is limited, such as in forensic or archival cases.
The library prep protocol you choose also sets constraints on which sequencing platform you can use and how your data will behave downstream. Short-read libraries for Illumina machines are structurally different from the circularized templates used in PacBio sequencing or the natively unmodified strands fed through nanopores. Choosing a library strategy is really choosing a whole data profile.
How the Major Sequencing Platforms Work
Three technology families dominate whole genome sequencing today, each with a fundamentally different approach to reading DNA.
Illumina’s platforms use sequencing by synthesis. Prepared library fragments are spread across a flow cell, amplified into clusters, and then read one base at a time as fluorescently labeled nucleotides are incorporated. The NovaSeq 6000, for example, can generate up to 6 terabases of data per run using this approach.5PubMed Central. TopoQual polishes circular consensus sequencing data and accurately predicts quality scores – Section: Abstract Illumina reads are short, typically 150 base pairs, but extremely accurate and produced at enormous scale. This makes the platform the workhorse for most clinical and population-scale sequencing projects.
PacBio’s HiFi sequencing takes a different route. DNA molecules are circularized, and a polymerase reads around the circle multiple times. The consensus of those repeated passes yields long reads averaging around 13.5 kilobases with accuracy around 99.8%.6Nature Biotechnology. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome That combination of length and accuracy is especially valuable for resolving repetitive regions and structural variants that short reads simply cannot span.
Oxford Nanopore Technology measures changes in electrical current as a single strand of DNA passes through a nanoscale pore in a membrane. An enzyme ratchets the DNA through the pore one base at a time, and shifts in the current are decoded into sequence by an algorithm.7Nature Biotechnology. Three decades of nanopore sequencing Nanopore’s standout feature is read length: individual reads can span tens or even hundreds of kilobases. Raw accuracy has historically been lower than Illumina or PacBio HiFi, but improvements in chemistry and computational base calling have steadily closed that gap.
From Raw Signal to Readable Bases
No sequencing platform directly hands you a clean DNA sequence. The raw output is either images of fluorescent clusters (Illumina), kinetic traces of polymerase activity (PacBio), or electrical current squiggles (Nanopore). Specialized software called base callers translates these signals into nucleotide sequences along with a per-base quality score. These quality scores follow the Phred scale, where Q10 means 90% accuracy, Q20 means 99%, and Q30 means 99.9%.8PubMed Central. Performance of neural network basecalling tools for Oxford Nanopore sequencing Modern base callers increasingly use neural networks, and their accuracy has improved considerably over the past few years, especially for nanopore data.
After base calling, the reads go through quality trimming and filtering. Low-quality bases at the ends of reads, adapter contamination, and reads that are too short to be useful are stripped out. This cleanup step is routine but has real effects on the reliability of everything downstream.
Aligning Reads to a Reference Genome
With clean reads in hand, the next step is figuring out where each one belongs in the genome. Alignment (or mapping) software takes each read and finds the best-matching position in a reference genome. Minimap2 is one of the most widely used aligners because it handles both short and long reads, working with accurate short reads of 100 base pairs or more as well as noisy long reads at error rates around 15%.9PubMed Central. Minimap2: pairwise alignment for nucleotide sequences Its ability to perform split-read alignment, where a single read maps to two different positions because of a structural rearrangement, is especially important for detecting large-scale genomic changes.
Alignment quality depends heavily on the reference genome itself. For years, a single linear reference (GRCh38) served as the standard for human sequencing. But a single reference cannot represent the diversity of human populations. The Human Pangenome Reference Consortium has been building a pangenome graph that stores a representative set of diverse haplotypes and their alignment, removing mapping biases inherent in a single linear reference.10Nature. A draft human pangenome reference Graph-based references built directly from genome assemblies can better represent complex multiallelic structural variants that a linear reference misses entirely.11Nature Biotechnology. Pangenome graph construction from genome alignments with Minigraph-Cactus Even with the current linear reference, modifying the aligner’s index to incorporate known population-specific variants has been shown to decrease false-negative variants by more than 9,500 and false positives by more than 7,000 in benchmark data.12PubMed Central. Enhancing SNV identification in whole-genome sequencing data through the incorporation of known genetic variants into the minimap2 index
Calling Variants
Once reads are mapped, variant calling software scans the alignments to identify positions where the sequenced individual differs from the reference. For single-nucleotide variants and small insertions or deletions, several callers compete for the top spot. In systematic benchmarks using gold-standard datasets from the Genome in a Bottle consortium, DeepVariant consistently showed the best performance and highest robustness among tested callers.13PubMed Central. Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery DRAGEN, a hardware-accelerated pipeline, performs similarly well, with no significant differences in its accuracy score compared to DeepVariant for single-nucleotide and small indel calling.14Scientific Reports. Accuracy and efficiency of germline variant calling pipelines for human genome data
When comparing GATK, the long-standing industry default, against DeepVariant in a trio sequencing study, DeepVariant produced a higher transition-to-transversion ratio (2.38 vs. 2.04), which is an indicator of higher-quality variant calls.15Scientific Reports. Comparison of GATK and DeepVariant by trio sequencing That said, GATK called more total variants overall, and the best choice of caller still depends somewhat on the data type and the downstream question. Many clinical labs run two callers and take the intersection or union to maximize confidence.
Structural Variant Detection
Structural variants, including deletions, insertions, duplications, inversions, and translocations of 50 base pairs or more, are among the hardest variant classes to detect reliably. Short-read sequencing has decent sensitivity for deletions (about 86% in one benchmark) but performs poorly for insertions (about 22% sensitivity).16PubMed Central. A Comparison of Structural Variant Calling from Short-Read and Nanopore-Based Whole-Genome Sequencing Using Optical Genome Mapping as a Benchmark Long-read nanopore sequencing with updated callers like Sniffles2 substantially outperforms short reads for most structural variant types, achieving 90% sensitivity for deletions and 74% for insertions in the same benchmark.
The gap between short and long reads is especially pronounced in repetitive regions. Short-read-based structural variant detection has significantly lower recall in repetitive sequences, particularly for small to intermediate-sized variants, compared to long-read approaches.17Human Genome Variation. Comparative evaluation of SNVs, indels, and structural variations detected with short- and long-read sequencing data In non-repetitive regions, the two technologies perform more similarly. For clinical labs, this means that a short-read-only pipeline may miss clinically relevant structural variants, especially insertions and events in repetitive DNA.
How Much Sequencing Is Enough
Coverage depth, meaning how many times each position in the genome is read on average, is one of the most important quality metrics in any WGS experiment. Sensitivity for detecting single-nucleotide and small indel variants increases with depth and reaches a plateau at about 40x mean depth. At that level, the sensitivity for both homozygous and heterozygous single-nucleotide variants exceeds 99.25%, and the breadth of coverage across disease-associated genes also plateaus.18PubMed Central. Characterizing sensitivity and coverage of clinical WGS as a diagnostic test for genetic disorders Most clinical WGS protocols therefore target 30x to 40x depth for germline analysis. Tumor sequencing often goes higher, sometimes to 60x or 100x, to catch variants present in only a fraction of tumor cells.
Breadth of coverage, the percentage of target bases sequenced at least a given number of times, matters as much as average depth. A genome with 40x average depth but large uncovered gaps is less useful than one with 35x and even distribution. This is why the fragmentation and library prep choices discussed earlier have real downstream consequences for diagnostic sensitivity.19Nature Reviews Genetics. Sequencing depth and coverage: key considerations in genomic analyses
Interpreting Variants in a Clinical Setting
Finding variants is only half the challenge. Deciding which ones matter clinically is where much of the intellectual work happens. The 2015 ACMG-AMP guidelines provide a standardized framework for classifying sequence variants into five tiers, from pathogenic to benign, using criteria that draw on population frequency, computational predictions, functional studies, and segregation data.20PubMed Central. Standards and Guidelines for the Interpretation of Sequence Variants: A Joint Consensus Recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology Automated tools like InterVar can apply 18 of the 28 ACMG criteria computationally, but the remaining 10 require manual expert input, including evidence from family studies and functional assays.21The American Journal of Human Genetics. InterVar: Clinical Interpretation of Genetic Variants by the 2015 ACMG-AMP Guidelines
For copy-number variants, a separate set of 2019 ACMG guidelines applies. Tools like ClassifyCNV implement these guidelines and match the pathogenicity category for about 81% of manually evaluated variants, with the remaining cases classified as uncertain and requiring further clinical review.22Scientific Reports. ClassifyCNV: a tool for clinical annotation of copy-number variants The persistent gap between automated and manual classification is one reason variant interpretation remains a bottleneck in clinical genomics. It takes trained geneticists and genetic counselors to bridge that gap, and the volume of variants per genome, typically several million, makes full manual review impractical without computational pre-filtering.
WGS Versus Exome Sequencing for Rare Disease Diagnosis
A recurring question in clinical genetics is whether whole genome sequencing adds enough diagnostic value over the cheaper alternative of exome sequencing, which targets only the roughly 2% of the genome that codes for proteins. A meta-analysis of pediatric rare disease studies found the pooled diagnostic yield was about 31% for genome sequencing versus 23% for exome sequencing, with roughly 1.7 times the odds of diagnosis for genome sequencing, though the difference did not quite reach conventional statistical significance.23Genetics in Medicine. A meta-analysis of diagnostic yield and clinical utility of genome and exome sequencing in pediatric rare and undiagnosed genetic diseases
Where does the extra yield come from? A large study sequencing 822 families found a molecular diagnosis in about 29% of them, and of those diagnosed, 28% had variants that specifically required genome sequencing for identification, including intronic variants, structural variants, copy-neutral inversions, and tandem repeat expansions.24PubMed. Genome Sequencing for Diagnosing Rare Diseases Interestingly, most families diagnosed after prior nondiagnostic exome sequencing had variants that could have been found by reanalysis of existing exome data or by adding copy-number variant calling to the exome pipeline. A Korean study of over 1,400 families found that about 15% of diagnosed families required genome sequencing specifically to detect complex structural variants, inversions, and noncoding variants that exome sequencing could not catch.25npj Genomic Medicine. Clinical utility of genome sequencing in rare diseases: lessons from a single-center study of 1,452 Korean families The practical takeaway is that WGS matters most when the diagnosis involves variant types that live outside exome territory.
Managing the Data Deluge
A single 30x human genome generates on the order of 60 to 70 gigabytes of compressed FASTQ data. Multiply that by thousands of samples in a biobank or clinical pipeline, and storage becomes a genuine logistical challenge. CRAM format, which compresses aligned reads against a reference genome, achieves 40–70% compression depending on the sequencing platform, translating the data into a much more manageable size without altering variant calls.26bioRxiv. CRAM compression: practical across-technologies considerations for large-scale sequencing projects Specialized tools push further: Genozip achieved roughly 6-fold compression on standard FASTQ files in benchmark testing, substantially outperforming older approaches, though at the cost of higher memory usage and run time.27Scientific Reports. A benchmark study of compression software for human short-read sequence data
Beyond compression, cloud computing has reshaped how large-scale sequencing projects manage data. Rather than moving terabytes between institutions, many consortia keep data in centralized cloud buckets and bring the analysis to the data. This shift introduces its own trade-offs in cost, reproducibility, and security, but it has made population-scale projects far more practical than they were even five years ago.
Cancer Sequencing Adds Layers of Complexity
WGS for cancer differs from germline sequencing in several important ways. Tumor samples are a heterogeneous mix of cancer cells and normal tissue, so variants may appear at much lower frequencies than the clean 50% or 100% you expect from a heterozygous or homozygous germline variant. A systematic evaluation of somatic mutation detection across six centers found that factors like tumor purity, input amount, library construction protocol, and bioinformatics pipeline all affected detection reproducibility and accuracy.28Nature Biotechnology. Toward best practice in cancer mutation detection with whole-genome and whole-exome sequencing Cancer WGS typically sequences a paired tumor and normal sample from the same patient, subtracting the germline background to isolate somatic mutations. The analytical challenge is greater, and the required depth is higher, but the payoff is a comprehensive view of the tumor’s mutational landscape, including structural rearrangements and mutational signatures that targeted panels miss.
Reading Epigenetic Marks Without Extra Steps
One of the more exciting recent developments is the ability to read DNA methylation, an epigenetic modification, directly during sequencing. Nanopore technology senses methylated bases as the DNA strand passes through the pore, eliminating the need for bisulfite conversion, which degrades DNA and introduces its own biases. Long nanopore reads allow researchers to examine methylation patterns across large genomic regions within single molecules, revealing heterogeneity that bulk-level methods underestimate.29PLOS Genetics. Genome-wide single-molecule analysis of long-read DNA methylation reveals heterogeneous patterns at heterochromatin that reflect nucleosome organisation This capability is still maturing, but it points toward a future where a single sequencing run yields both the genome sequence and a map of its epigenetic regulation.
Privacy Risks in Genomic Data
A whole genome is the ultimate personal identifier. Privacy attacks on genomic data fall into several categories, including identity tracing, attribute disclosure (inferring traits or disease status), and completion attacks that fill in missing genotype information from partial data.30Briefings in Bioinformatics. Ensuring privacy and security of genomic data and functionalities Membership inference attacks, which try to determine whether a specific individual was part of a research cohort, can succeed with high power when more than 250 variants are shared from the case group.31bioRxiv. Safeguarding Privacy in Genome Research: A Comprehensive Framework for Authors Emerging solutions include hybrid systems combining blockchain for access control with homomorphic encryption that allows computation on encrypted data without ever exposing raw sequences.32Ethics, Medicine and Public Health. Blockchain and homomorphic encryption for genomic and health data sharing: An ethical perspective For anyone generating or sharing WGS data, understanding these risks is no longer optional. Consent frameworks, data-use agreements, and technical safeguards all need to keep pace with the growing ease of sequencing and the permanence of genomic information.