The Telomere-to-Telomere (T2T) Consortium published the first truly complete sequence of a human genome in 2022, a gapless assembly spanning 3.055 billion base pairs that finally addressed the roughly 8 percent of the genome left unfinished by the Human Genome Project two decades earlier.1Science. The complete sequence of a human genome That missing fraction was not random filler. It contained centromeres, segmental duplications, the short arms of five chromosomes, and other structurally complex regions that older sequencing methods could not read through. Filling those gaps turned up almost 200 million new base pairs and nearly 2,000 predicted genes, reshaping how researchers study everything from chromosome biology to human disease.
What Was Actually Missing
When the Human Genome Project declared victory in 2003, the celebratory language overshadowed a practical reality: the reference genome, eventually refined into a version called GRCh38, still contained hundreds of gaps. These gaps were concentrated in regions loaded with highly repetitive DNA, where short sequencing reads could not be placed with any confidence. The prior reference contained over 270 million bases in segmental duplications and low-mappability zones across the autosomes alone, even when modeled centromeres were included.2Cell Genomics. Expanded benchmark for 7 genomes from the Genome in a Bottle Consortium – Section: Results Centromeres, built from vast tandem arrays of satellite DNA, were essentially represented by placeholder models rather than real sequence. The short arms of the five acrocentric chromosomes (13, 14, 15, 21, and 22), which house ribosomal RNA gene arrays critical for protein production, were also left unresolved.
These were not minor cosmetic gaps. Centromeres are the structures that ensure chromosomes split correctly during cell division. Segmental duplications are hotspots for gene birth and structural rearrangement. Without their actual sequences, researchers could not study variation in these regions, could not map sequencing reads to them, and could not assess their role in disease. The 8 percent figure sounds modest until you realize it represented the most biologically active and evolutionarily dynamic parts of the genome.
Why a Hydatidiform Mole
One of the cleverest decisions the T2T project made was its choice of source material. A normal human cell carries two copies of each chromosome, one from each parent, and those copies differ at millions of positions. Assembling a genome means untangling which reads belong to which copy, a computationally brutal task in repetitive regions. The consortium sidestepped this by sequencing CHM13, a cell line derived from a complete hydatidiform mole. This is a type of abnormal pregnancy in which the resulting tissue carries two identical copies of the father’s genome and none of the mother’s. The genome is essentially homozygous, meaning both chromosome copies are nearly identical, which dramatically simplifies assembly. The T2T-CHM13 reference includes gapless assemblies for all 22 autosomes plus the X chromosome.1Science. The complete sequence of a human genome Because the mole lacks a Y chromosome, that piece had to wait.
The Sequencing Technologies That Made It Possible
Two long-read sequencing platforms converged to crack the problem. PacBio’s HiFi (High-Fidelity) sequencing produces reads typically 10 to 25 kilobases long with a median accuracy of 99.9 percent and reliable resolution of tricky homopolymer stretches (runs of the same base repeated five or more times).3PubMed Central. Long and Accurate: How HiFi Sequencing is Transforming Genomics Those reads are long enough to span many repetitive elements and accurate enough to distinguish near-identical copies.
Oxford Nanopore’s ultra-long read technology fills a different niche. These reads can stretch beyond 100 kilobases, sometimes reaching a megabase. Earlier work had shown that adding even modest coverage of ultra-long nanopore reads could more than double assembly contiguity and resolve complex loci like the major histocompatibility complex in a single piece.4PubMed Central. Nanopore sequencing and assembly of a human genome with ultra-long reads Ultra-long reads also allowed direct measurement of telomere repeat length, something short-read data simply cannot do.
Neither technology alone was sufficient. HiFi reads provided the base-level accuracy needed to build an initial graph of how sequence segments connect. Ultra-long nanopore reads then bridged the remaining tangles, spanning entire satellite arrays that even 20-kilobase reads could not cross. The assembly tool Verkko, developed to handle T2T-scale projects, formalizes this combination: it starts by building a graph from HiFi reads and then progressively simplifies the graph using ultra-long reads and haplotype-specific markers.5PubMed Central. Telomere-to-telomere assembly of diploid chromosomes with Verkko – Section: Results The result is an iterative, automated pipeline capable of producing complete diploid genome assemblies, not just the homozygous mole genome that launched the project.
What the Completed Genome Revealed
The T2T-CHM13 assembly added nearly 200 million base pairs to the human reference. Within that new sequence, the consortium identified 1,956 gene predictions, 99 of which are predicted to encode proteins.1Science. The complete sequence of a human genome That is not a trivial haul. For context, the entire human genome contains roughly 20,000 protein-coding genes, so the newly resolved regions harbor a meaningful fraction of genes that had never been placed in a reference before.
The completed regions also included the first full views of all human centromeric satellite arrays. These arrays are built from tandem repeats that can span millions of bases on each chromosome, and their sequences vary in ways no one had been able to study systematically. Additionally, the project resolved recent segmental duplications, identifying 182 candidate protein-encoding genes as well as complete sequences for gene models that had been structurally ambiguous in prior assemblies.6PubMed Central. Segmental duplications and their variation in a complete human genome – Section: Results And the short arms of all five acrocentric chromosomes, home to the ribosomal DNA arrays that cells depend on for manufacturing ribosomes, were sequenced end to end for the first time.7PubMed Central. Short arms of human acrocentric chromosomes and the completion of the human genome sequence
Completing the Y Chromosome
Because the CHM13 cell line lacks a Y chromosome, the consortium tackled it separately, publishing the complete Y sequence in 2023. The finished Y chromosome clocks in at 62,460,029 base pairs, adding over 30 million base pairs of sequence that had been missing or misassembled in GRCh38.8Nature. The complete sequence of a human Y chromosome This is staggering: more than half of the Y chromosome’s sequence had been absent from the reference genome that researchers worldwide had been using.
The newly resolved Y sequence revealed 41 additional protein-coding genes, most belonging to the TSPY gene family, and showed the complete structures of gene families like DAZ and RBMY that are critical for male fertility.9PubMed Central. The complete sequence of a human Y chromosome Perhaps the most visually striking finding was the organization of the heterochromatic Yq12 region, a vast stretch of the Y chromosome’s long arm that turned out to contain alternating blocks of human satellite 1 and satellite 3 sequences. Before the T2T assembly, this region was essentially a black box.
Centromeres Up Close
With complete centromeric sequences in hand, researchers could finally study these regions at the level of individual DNA bases. Centromeres matter for a straightforward reason: if they malfunction, chromosomes mis-segregate during cell division, leading to conditions like Down syndrome or driving cancer progression. Yet before T2T, no one had their full sequences.
The complete assembly revealed extensive variation between the two copies (haplotypes) of each centromere in a diploid genome, including differences in array size, internal repeat organization, and the evolutionary layering of related satellite repeats within arrays.10PubMed Central. Haplotype-resolved centromeric chromatin organization from a complete diploid human genome – Section: Haplotype-resolved annotation reveals multi-scale variation in human centromeres Researchers have also begun mapping the epigenetic landscape of centromeres on these complete assemblies, tracking where centromere-specific proteins bind and how methylation patterns vary across satellite arrays.11Cell Genomics. Telomere-to-telomere mapping of chromatin architecture and epigenetic plasticity at human centromeres – Section: Results This work is uncovering how centromere identity is maintained and how it can drift between haplotypes, questions that were unanswerable when the sequences themselves were unknown.
Non-Canonical DNA Structures
DNA does not always twist into the classic double helix. Under certain sequence conditions, it can fold into alternative shapes: G-quadruplexes (stacked guanine tetrads), Z-DNA (left-handed helices), triplexes, and other conformations collectively called non-B DNA. These structures influence gene regulation, replication timing, and genome instability, but studying their distribution requires knowing the full underlying sequence.
Analyses of the T2T assembly found an overrepresentation of most non-B DNA motif types in the newly added sequences compared to the older GRCh38 reference.12Nucleic Acids Research. Non-canonical DNA in human and other ape telomere-to-telomere genomes – Section: Results In other words, the parts of the genome that had been missing were disproportionately rich in sequences capable of forming unusual structures. Fully resolved genomes now reveal the true abundance and chromosomal distribution of these motifs, including within satellite and other repetitive regions that were previously inaccessible.13PubMed. Unraveling Non-B DNA Structures in the Era of Telomere-to-Telomere Genomes This matters because non-B DNA is increasingly linked to mutational hotspots and structural rearrangements. Understanding where these motifs concentrate gives researchers new leads on why certain genomic regions are prone to breaks and rearrangements.
Cleaning Up Variant Calling
A genome reference is not just a curiosity. It is the coordinate system that every clinical and research sequencing pipeline uses. When a patient’s genome is sequenced, the reads are mapped against the reference to find differences, called variants. If the reference itself has errors or gaps, the mapping goes wrong in predictable ways: real variants get missed, and phantom variants appear where none exist.
Switching from GRCh38 to T2T-CHM13 improved read mapping and variant calling across the board, tested on more than 3,200 globally diverse samples sequenced with short reads and additional samples sequenced with long reads.14PubMed Central. A complete reference genome improves analysis of human genetic variation The improvements went in both directions: researchers identified hundreds of thousands of variants per sample in regions that had been invisible before, and simultaneously eliminated tens of thousands of spurious variant calls per sample. In 269 medically relevant genes, false-positive variant calls dropped by up to a factor of 12. For clinical genomics, where a single false call can trigger unnecessary follow-up or a missed call can mean a delayed diagnosis, that kind of cleanup is transformative.
The complete reference also opens the door to studying complex structural variants, rearrangements involving multiple breakpoints that are important contributors to rare genetic diseases but were poorly characterized when the reference itself was incomplete.15European Journal of Human Genetics. Rare disease genomics in an era of human pangenomics and telomere-to-telomere genome references – Section: A telomere-to-telomere reference improves analysis of genomic variation, particularly complex SVs With the repetitive regions now resolved, pipelines can characterize deletions, duplications, inversions, and translocations that previously fell in unmappable territory.
From One Genome to a Pangenome
One person’s genome, however complete, cannot capture the full range of human genetic diversity. The T2T-CHM13 assembly is a single haploid sequence from a cell line of European ancestry. Entire structural variants common in other populations are simply absent from it. Recognizing this, the Human Pangenome Reference Consortium (HPRC) set out to build a reference that represents many genomes at once.
The HPRC’s initial effort assembled high-quality diploid genomes from 47 individuals (94 haplotypes), drawn from diverse global populations, using the trio-based graph-partitioning approach of the assembler hifiasm.16Nature. Semi-automated assembly of high-quality diploid human reference genomes – Section: A look towards the future The goal is a graph-based reference in which each path through the graph represents a real human haplotype, so that variant calling does not force every person’s genome into the shape of a single template.
The T2T technologies were essential scaffolding for this work. Generating complete, haplotype-phased assemblies for dozens of individuals requires the same long-read strategies and assembly algorithms developed for T2T-CHM13, now scaled up. The Verkko assembler, originally built for the mole’s homozygous genome, was specifically extended to handle diploid genomes with two divergent haplotypes. The pangenome project is, in many ways, the T2T approach applied at population scale.
Ape Genomes and Evolutionary Context
The same sequencing and assembly methods have been turned on our closest relatives. The T2T Consortium and allied groups have now produced complete or near-complete genome assemblies for multiple great ape species. These assemblies revealed that genetic divergence between ape species is substantially greater than older, incomplete references had suggested. Across complete ape genomes, 12.5 to 27.3 percent of an ape genome failed to align to the human reference or was inconsistent with simple one-to-one alignment, driven largely by rapidly evolving repetitive and structurally variable regions.17Nature. Complete sequencing of ape genomes – Section: Results
That divergence was invisible when incomplete genomes were compared to an incomplete human reference, because the most variable regions were missing from both sides of the comparison. Complete assemblies now allow researchers to study lineage-specific segmental duplications, centromeric DNA evolution, and subterminal heterochromatin without the bias introduced by mapping everything to a human template.18PubMed Central. Complete sequencing of ape genomes For understanding how species diverge at the structural level, rather than just at individual nucleotide positions, this is a fundamentally different kind of data.
What the Reference Still Cannot Do
For all its achievements, a single reference genome has inherent limitations that are worth being clear-eyed about. T2T-CHM13 represents one person’s genome, and a somewhat unusual one at that: it comes from an abnormal tissue with no heterozygosity. Structural variants that exist only in other populations, rare alleles, and the combinatorial complexity of having two different haplotypes in every cell are all beyond its scope as a lone reference. The pangenome effort addresses this, but it is still in its early phases.
There is also the question of functional annotation. Having the sequence of a gene does not automatically tell you what it does. The 99 newly predicted protein-coding genes in the T2T assembly still need experimental validation: do they produce functional proteins, in which tissues, under what conditions? The centromeric satellite arrays are now sequenced, but their functional elements, particularly the epigenetic marks that define where centromere proteins actually bind, are only beginning to be mapped across individuals. The sequence is the foundation, not the finished building.
Clinical adoption is another practical bottleneck. Most sequencing pipelines worldwide still use GRCh38 as their default coordinate system. Switching references means re-running analyses, updating databases, retraining clinical variant-interpretation workflows, and reconciling years of data annotated against the old coordinates. The benefits are clear, especially the reduction in false positives in medically relevant genes, but the transition is logistically complex and will take time.
Epigenetic Layers on a Complete Map
Sequence alone is a static blueprint. Cells regulate which parts of the genome are active through chemical modifications to DNA and to the histone proteins that package it. Methylation of cytosine bases, for instance, typically silences nearby genes, while specific histone marks signal whether a region is transcriptionally active, compacted into silent heterochromatin, or functioning as a centromere. Before T2T, studying these marks in repetitive regions was nearly impossible because the underlying sequences could not be uniquely identified.
Nanopore sequencing has a native advantage here: as DNA passes through the pore, the electrical signal encodes not just the base sequence but also chemical modifications like methylation. Researchers have used this property to map centromere-specific proteins (CENP-A), repressive chromatin marks (H3K9me3), and CpG methylation directly onto the complete T2T assembly using techniques that combine directed methylation with long-read sequencing.11Cell Genomics. Telomere-to-telomere mapping of chromatin architecture and epigenetic plasticity at human centromeres – Section: Results The resulting maps show, for the first time at single-molecule resolution, how epigenetic states are distributed across centromeric satellite arrays on individual chromosomes. Early findings suggest that the active centromere domain, where CENP-A sits, occupies a surprisingly narrow window within the much larger satellite array, and that this window can shift position between haplotypes. Understanding what drives that positioning could eventually shed light on chromosome segregation errors, including those that cause aneuploidy in cancer and age-related fertility decline.