The human pangenome is a new kind of reference genome built not from one person’s DNA but from hundreds of diverse individuals, organized as a graph structure that captures the full range of human genetic variation. It replaces the single linear sequence that scientists have relied on since the Human Genome Project with something far more representative of our species. The shift matters because the old reference, drawn largely from a single individual of mixed European and African ancestry, missed millions of base pairs of sequence and entire categories of variation that exist across global populations, creating blind spots in medical research and clinical diagnostics.
Why a Single Reference Genome Was Not Enough
The original human reference genome was a landmark achievement, but it was always a compromise. It represented one version of the human genome assembled from a handful of donors, with one individual contributing roughly 70% of the final sequence. For years, that reference contained hundreds of gaps in hard-to-sequence repetitive regions, and it offered no way to represent the natural genetic differences between people. When researchers compared a patient’s DNA to this single reference, any variant that happened to differ from that one person’s genome looked like a deviation, even if it was perfectly common in the patient’s own population.
Two advances have recently addressed these shortcomings. The Telomere-to-Telomere (T2T) Consortium produced the first truly complete, gap-free human genome sequence, adding roughly 200 million base pairs of previously unresolved sequence and correcting thousands of structural errors in the old reference.1Science. A complete reference genome improves analysis of human genetic variation That complete genome eliminated tens of thousands of false-positive variant calls per sample and reduced errors in medically relevant genes by up to twelve-fold. But even a perfect linear genome from one person cannot represent the diversity of nearly eight billion people. That is where the pangenome comes in.
The Human Pangenome Reference Consortium (HPRC) was formed with the goal of creating high-quality diploid genome assemblies from globally diverse individuals, capturing variation from single-letter changes all the way up to large structural rearrangements.2PubMed Central. The Human Pangenome Project: a global resource to map genomic diversity Rather than replacing the old reference with a slightly better one, the consortium set out to replace the entire concept of a single linear reference.
How the Graph-Based Pangenome Works
If you imagine a traditional reference genome as a single railroad track running from one end of a chromosome to the other, a pangenome graph is more like a highway system with alternate routes. Where people’s genomes agree, there is one shared path. Where they differ, the graph branches into alternate paths representing insertions, deletions, duplications, or inversions, then merges back together. These branching structures, called variation graphs, compactly encode genetic variation across a population, including large-scale rearrangements that a simple linear sequence cannot represent at all.3PubMed Central. Variation graph toolkit improves read mapping by representing genetic variation in the reference
Building this graph required a new generation of sequencing and assembly methods. The HPRC’s initial work showed that combining highly accurate long-read sequencing with parent-child data and graph-based phasing produced the most complete diploid assemblies, with only about four gaps per chromosome on average.4Nature. Semi-automated assembly of high-quality diploid human reference genomes “Diploid” is the key word: unlike the old reference, which collapsed both copies of each chromosome into one, the pangenome keeps both copies of every person’s chromosomes separate, preserving the combinations of variants that actually travel together on the same strand of DNA.
The practical payoff is that when you map a patient’s sequencing reads against this graph instead of a single linear sequence, the reads land more accurately because the graph already contains many of the variants the patient carries. A fast read-mapping tool called Giraffe can align both short and long reads to a pangenome graph containing data from more than 450 human haplotypes at speeds comparable to traditional linear mappers, while producing similar or improved variant-calling results.5PubMed Central. Rapid, accurate long- and short-read mapping to large pangenome graphs with vg Giraffe
Structural Variants Are the Big Payoff
Most conversation about human genetic variation focuses on single-letter DNA changes, and for good reason: they make up about 99% of the differences between any two genomes. But the pangenome has revealed that when you count how many actual base pairs are affected, structural variants (insertions, deletions, duplications, and inversions of 50 or more base pairs) account for roughly 88% of the variant sequence between two people.6medRxiv. A high-resolution human pangenome structural variant resource for improved disease association These large rearrangements are the variants most likely to disrupt genes, alter regulatory regions, or change how chromosomes fold. Yet the old linear reference was poorly suited to detecting them because the sequencing reads that span a structural variant often fail to align properly to a genome that does not contain it.
The pangenome graph sidesteps that problem. Because the graph already contains known structural variants from hundreds of individuals, reads that carry those variants align cleanly rather than getting misaligned or thrown away. This is not an incremental improvement. Research on gastric cancer patients found that aligning to a pangenome graph detected roughly 65% more structural variants than aligning to the old linear reference, with the majority of newly found variants invisible to linear methods entirely.7Life Science Alliance. Gastric cancer genomics study using reference human pangenomes
Sequences That Were Simply Missing
Beyond structural variants, the pangenome has uncovered tens of megabases of DNA sequence that were absent from previous reference genomes altogether. One large-scale analysis identified over 45,000 non-redundant sequences totaling about 60 megabases that do not appear in the standard reference.8Nucleic Acids Research. Human pangenome analysis of sequences missing from the reference genome reveals their widespread evolutionary, phenotypic, and functional roles About 78% of these are repetitive sequences, which partly explains why earlier short-read technologies missed them. But a meaningful fraction falls in or near genes, carrying potential functional consequences that were invisible to any analysis built on the old reference.
Many of these missing sequences show population-specific patterns. In the study above, over 12,000 novel sequences were found in East Asian populations, with nearly 5,000 specific to that group. The old reference’s bias toward a small number of donors meant that sequences common in some populations simply did not exist in the reference that clinicians and researchers worldwide were using. Discovering these sequences is not just an academic exercise; some sit near genes implicated in disease, immune function, or drug metabolism.
Diagnosing Rare Genetic Diseases
One of the most immediate clinical applications of the pangenome is in rare disease diagnostics. Roughly a third of patients with suspected genetic conditions go undiagnosed after standard testing, and a significant fraction of those unsolved cases involve structural variants that conventional pipelines miss. Using a pangenome graph to analyze genomes from rare disease patients has already yielded new diagnostic findings.
In one case, pangenome-based analysis uncovered a 14.5-kilobase deletion in a gene called KMT2E in a child with low muscle tone, an unusually large head, and developmental delay. The deletion, which spanned several exons and was predicted to cause loss of function, had been missed by prior testing. The child’s symptoms matched the clinical profile already described for harmful variants in that gene, providing a clear diagnosis.9Nature Communications. Pangenome graphs improve the analysis of structural variants in rare genetic diseases That deletion sat in a region that the old linear reference and standard variant-calling tools struggled with, making it effectively invisible until the pangenome approach was applied.
Adopting pangenome methods in clinical settings is not yet routine, and doing so will require careful selection of diverse reference individuals, community engagement around consent and data sharing, extensive data generation using multiple sequencing technologies, and specialized computational expertise.10European Journal of Human Genetics. Rare disease genomics in an era of human pangenomics and telomere-to-telomere genome references But the diagnostic yield improvements are real enough that clinical labs are beginning to evaluate these tools seriously.
Sharper Cancer Mutation Detection
Cancer genomes are messy. Tumors accumulate structural variants at a high rate, and distinguishing true cancer-specific mutations from false positives caused by reference bias has been a persistent headache. When sequencing reads from a tumor are aligned to a linear reference that does not contain the patient’s normal germline variants, those germline variants get flagged as somatic mutations, cluttering the results with noise.
Pangenome-based approaches reduce that noise substantially. A tool called SVPG, which aligns tumor and matched normal tissue reads to a pangenome graph, achieved markedly higher accuracy on benchmark cancer datasets than traditional linear-reference callers, with far fewer false positives.11Nature Methods. SVPG: a pangenome-based structural variant detection approach and rapid augmentation of pangenome graphs with new samples A separate study developed a filtering method that combines pangenome alignment with de novo assembly of the patient’s germline genome, dramatically reducing false-positive somatic structural variant calls across cancer cell lines while preserving sensitivity.12PubMed. Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly
The practical upside for oncology is that fewer false positives means less wasted time chasing artifacts, more confidence in the mutations that are called, and ultimately more reliable information for guiding treatment. As long-read sequencing becomes cheaper, pangenome-based cancer genomics is likely to move from benchmarking studies into clinical pipelines.
Closing the Diversity Gap
The equity dimension of the pangenome project is hard to overstate. Africa harbors more genetic variation than the rest of the world combined, yet roughly 1% of genomes in major databases come from individuals of African ancestry. As one recent review put it bluntly, this is not merely an equity problem but a scientific error that distorts drug dosing, degrades risk scores, and undermines precision medicine globally.13Trends in Genetics. African genomes, pangenomes, and precision medicine
Polygenic risk scores, which estimate a person’s genetic likelihood of developing conditions like heart disease or diabetes, perform best in the populations they were trained on and lose accuracy in others. Since most large genetic studies have been conducted in European-ancestry cohorts, risk scores derived from those studies work poorly for people of African, South Asian, or Indigenous American descent. A pangenome that genuinely represents global diversity would improve variant calling and association studies across populations, helping close this performance gap.
Population-specific pangenome efforts are already emerging. A Vietnamese pangenome resource, for example, built a population-specific reference that enabled more accurate genetic imputation for Vietnamese individuals than any existing reference panel, with particular improvements in imputing immune-system genes that vary widely across populations.14Nature Communications. VN1K is a pangenome-informed multi-omics and phenomics resource for the Vietnamese population Similar efforts are underway in African, East Asian, and other underrepresented groups, each contributing both to their own populations’ medical needs and to the broader global pangenome.
Tracing Ancient DNA Through Modern Genomes
The pangenome has opened a surprising window into human evolution. When modern humans migrated out of Africa and encountered Neanderthals and Denisovans, interbreeding left fragments of archaic DNA scattered through non-African genomes. Detecting those fragments has been difficult for structural variants, because the old reference genome could not represent them properly. Pangenome methods have changed that.
A study integrating high-quality phased assemblies from Papua New Guinean individuals with 94 other diverse genomes mapped introgressed structural variants across modern human populations. The findings were striking: archaic structural variants were enriched in gene-containing regions, with 47% overlapping genes, and were most abundant in Papua New Guinean genomes. The study identified 11 centromeres likely derived from archaic hominins and pinpointed 16 adaptive structural variants, many linked to immune-related genes.15PubMed Central. A global map for introgressed structural variation and selection in humans Separately, the Chinese Pangenome Project characterized over 700 megabases of introgressed sequences and found an enrichment of Denisovan-like archaic segments in East Asian genomes.16Cell Genomics. The Human Pangenome: What It Is & Why It Matters
These are not just curiosities. Some archaic structural variants appear to have been positively selected because they conferred advantages, particularly in immune function. Understanding which archaic variants persist and where they sit in the genome helps explain present-day differences in disease susceptibility across populations and adds a new dimension to precision medicine that was inaccessible before pangenome methods existed.
The Computational Challenge of Replacing a Line with a Graph
Switching from a linear reference to a graph pangenome means rethinking virtually every tool in the genomics pipeline. Decades of software development assumed a single coordinate system: position 1 to position 3 billion along each chromosome. A graph has no single coordinate axis. Every aligner, variant caller, annotation database, and visualization tool that genomics labs depend on was designed for lines, not graphs.
On the practical side, the field still needs widely accepted file formats for sequence graphs and for storing read alignments against them.17PubMed Central. Computational pan-genomics: status, promises and challenges The text-based Graphical Fragment Assembly (GFA) format works for the graph structure itself but becomes unwieldy when it also needs to encode the paths of hundreds of individual genomes through the graph, since common compression tools cannot efficiently handle the redundancy of highly similar paths spanning megabases.18Oxford Academic. GBZ file format for pangenome graphs New compressed formats like GBZ have been developed to address this, but the ecosystem of tools that can read and write them is still maturing.
On the theoretical side, many fundamental algorithms in genomics have known computational properties when operating on linear strings. When those same problems are posed on graphs, the computational difficulty can change in ways that are not yet fully understood.19PubMed Central. Computational graph pangenomics: a tutorial on data structures and their applications Ambitious national sequencing projects aiming to sequence a million or more individuals will need pangenome infrastructure that can scale to that level, and that infrastructure does not fully exist yet. The biology is ahead of the engineering, which is a good problem to have but a real bottleneck for clinical adoption.
What the Pangenome Cannot Yet Do
For all its promise, the current pangenome has real limitations worth understanding. The HPRC’s initial releases are built from fewer than 100 individuals, which, while far better than one, still underrepresents many of the world’s populations. Expanding to the hundreds or thousands of high-quality diploid assemblies needed for a genuinely global reference is expensive and logistically complex. Each genome requires multiple sequencing technologies, parental samples for phasing, and substantial computational resources to assemble.
Ethical and governance questions are also evolving. Genome data from Indigenous and historically marginalized communities carries particular sensitivities around consent, data sovereignty, and benefit-sharing. The HPRC has committed to ethical frameworks that address community engagement and informed consent, but the specifics of data governance, such as who controls access and how benefits flow back to contributing communities, remain works in progress. Getting the science right without the ethics right would undermine the entire enterprise.
There is also a training gap. Most clinical geneticists and bioinformaticians were trained on linear-reference workflows. Transitioning to pangenome methods requires not just new software but new intuitions about how to interpret results in a graph context. A variant that looks straightforward on a linear reference can appear as a complex set of paths in a graph, and understanding what that means for a patient takes practice that most labs have not yet had.
How Personal Genomes Fit Into the Picture
One misconception about the pangenome is that it replaces the need for individual genome sequencing. It does not. The pangenome is a reference, a map that makes it easier to interpret any individual’s genome by providing a richer backdrop of known variation. You still need to sequence a patient’s DNA and compare it to the pangenome; the pangenome just makes that comparison more accurate and complete.
An emerging approach goes further: building a personal genome assembly for each individual using long-read sequencing, then using the pangenome graph as a scaffold to place that personal assembly in context. In cancer genomics, combining pangenome alignment with a de novo assembly of the patient’s germline genome produced the cleanest results for separating true tumor mutations from background noise. In rare disease settings, having both a patient’s personal assembly and the pangenome graph allowed detection of variants that neither approach would have found alone.
As sequencing costs continue to fall, the combination of a rich pangenome reference and high-quality personal assemblies may become the standard of care for genomic medicine. That future is probably years away for routine clinical use, but the research infrastructure is being laid now, and the early results from cancer, rare disease, and population genetics all point in the same direction: the single linear reference genome has served its purpose, and the field is moving on.