Tagmentation is a technique that uses an enzyme called a transposase to simultaneously cut DNA and attach short adapter sequences to the resulting fragments, combining what used to be two separate and time-consuming steps in preparing DNA for sequencing. The name itself is a portmanteau of “tagging” and “fragmentation,” which captures the essence of what happens in a single reaction tube. The method has become central to modern genomics because it can compress hours of bench work into minutes, but the details of how the enzyme actually accomplishes this feat, and where the technique trips up, are worth understanding for anyone working with or reading about sequencing data.
How Tn5 Transposase Does the Work
The workhorse enzyme behind tagmentation is a hyperactive form of the Tn5 transposase, a protein borrowed from a bacterial mobile genetic element. In nature, Tn5 is a transposon, a segment of DNA that can move itself from one location to another within a genome. Researchers engineered a version of this enzyme that is far more active than the wild-type form, and they load it with synthetic DNA adapters before it ever touches the sample. The resulting complex of enzyme plus adapters is called a transposome.
When the transposome encounters target DNA, the enzyme cuts both strands and simultaneously stitches the adapter sequences onto the newly exposed ends. The chemistry involves a series of precise molecular events. First, the transposase nicks one strand of the target DNA, freeing a reactive chemical group. That group then attacks the opposite strand, forming a temporary hairpin structure. The enzyme resolves the hairpin, freeing another reactive group that joins to the adapter DNA in a strand-transfer reaction.1Journal of Biological Chemistry. Hairpin Formation in Tn5 Transposition Because Tn5 works as a dimer, two copies of the enzyme collaborate to process both ends of the insertion site, resulting in a DNA fragment with adapter sequences on each end. The process leaves small gaps of nine base pairs at each insertion site, which are later filled in during a brief repair step.
What makes this so useful for sequencing is that the adapters contain all the sequences a sequencing machine needs to recognize and read the DNA fragment. In traditional library preparation, you would fragment the DNA mechanically (by sonication, for instance), repair the ragged ends, add short connector sequences, ligate the adapters, and then clean up the product. Tagmentation collapses all of that into one enzymatic step lasting about five minutes, followed by cleanup and a short round of amplification to enrich for properly tagged fragments.
Controlling Fragment Size
One of the practical questions anyone setting up a tagmentation reaction faces is how to control the size of the resulting DNA fragments. Fragment length matters because different sequencing applications require different insert sizes. Whole-genome sequencing benefits from a range of fragment lengths, while some targeted methods work best with tightly controlled, shorter fragments.
The main dial you can turn is enzyme concentration relative to the amount of input DNA. More transposase means more frequent insertions, which produces shorter fragments. Less enzyme means longer stretches of DNA survive between insertion events. Researchers have also engineered Tn5 variants with mutations in the DNA-binding domain that make this tuning even more precise. One such variant carries a point mutation called R27S, which allows finer adjustment of the fragment size distribution simply by changing the ratio of enzyme to DNA.2PubMed Central. Large-Scale Low-Cost NGS Library Preparation Using a Robust Tn5 Purification and Tagmentation Protocol Reaction temperature, incubation time, and the presence of molecular crowding agents also influence the outcome. Crowding agents, which mimic the packed conditions inside a cell, both modulate fragment lengths and allow the enzyme to work efficiently on extremely small amounts of starting material, down to sub-picogram quantities of complementary DNA.3PubMed Central. Tn5 transposase and tagmentation procedures for massively scaled sequencing projects
The GC Bias Problem
Tagmentation is not perfectly random. Tn5 has a well-documented preference for inserting into GC-rich DNA, meaning regions with more guanine and cytosine nucleotides receive more insertions than regions dominated by adenine and thymine. This is not a subtle effect. Studies across multiple species have found a consistent sequence motif at Tn5 insertion sites, with a characteristic G/C pair at the edges of the nine-base-pair core that the enzyme recognizes.4PubMed Central. Comprehensive understanding of Tn5 insertion preference improves transcription regulatory element identification The recognition sequence contains a palindromic pattern, which makes sense because the enzyme operates as a two-part dimer that interacts with the DNA from both sides.5NAR Genomics and Bioinformatics. Correction of transposase sequence bias in ATAC-seq data with rule ensemble modeling
The practical consequence is uneven coverage. Genomes that are naturally AT-rich, such as those of certain parasites and fungi, end up with poorer representation of AT-rich regions in the sequencing data. Comparisons between different transposases have shown that both Tn5 and the phage Mu transposase share this GC preference, leading to less uniform spacing of insertions in AT-rich genomes compared with other transposon systems.6PubMed Central. Insertion site preference of Mu, Tn5, and Tn7 transposons That said, only about 16 to 29 percent of Tn5 insertion sites in any given genome actually fall inside the predicted motif sites, which means the sequence preference is real but not the sole factor governing where the enzyme lands.4PubMed Central. Comprehensive understanding of Tn5 insertion preference improves transcription regulatory element identification Chromatin structure, DNA accessibility, and physical constraints of the DNA molecule also play roles.
This bias has prompted engineering efforts. A mutant transposase called Tn5-059 was developed through protein engineering specifically to reduce the GC insertion preference. It demonstrably lowers AT dropout and improves the uniformity of genome coverage across both bacterial and human genomes.7PubMed Central. Improved genome sequencing using an engineered transposase Computational correction methods have also been developed to model and remove the bias in downstream analysis, particularly for applications like ATAC-seq where accurate quantification of insertion frequency is the whole point of the experiment.5NAR Genomics and Bioinformatics. Correction of transposase sequence bias in ATAC-seq data with rule ensemble modeling
Mapping Open Chromatin with ATAC-seq
The application that probably did the most to make tagmentation famous outside of routine sequencing is ATAC-seq, which stands for Assay for Transposase-Accessible Chromatin with sequencing. The idea is elegant: if you expose intact cell nuclei to the Tn5 transposome, the enzyme can only insert adapters into regions of the genome where the DNA is physically accessible, not wrapped tightly around proteins. By sequencing the fragments that get tagged, you create a genome-wide map of which regions of chromatin are open and potentially active in gene regulation.8PubMed Central. ATAC-seq: A Method for Assaying Chromatin Accessibility Genome-Wide
ATAC-seq quickly became one of the most popular methods in epigenomics because it requires far fewer cells than older techniques for measuring chromatin accessibility, and the protocol is fast. The GC bias discussed above is a particular concern in ATAC-seq, though, because the whole experiment hinges on accurately interpreting where Tn5 did and did not insert. If the enzyme has an inherent preference for certain sequences, some “open” regions might be over-counted and others missed, which is why bias correction is an active area of research in this field.
Single-cell versions of ATAC-seq have pushed tagmentation further into high-throughput territory. A technique called single-cell combinatorial indexing uses Tn5 to tag chromatin in thousands of individual cells in a single experiment, assigning molecular barcodes at each step so that the resulting data can be traced back to its cell of origin. Newer chemistry, such as a uracil-based adapter switching approach called s3-ATAC, has improved the efficiency of this process, achieving a 6-to-13-fold increase in usable sequencing reads per cell compared with earlier methods when applied to mouse brain tissue.9PubMed Central. High-content single-cell combinatorial indexing
Tagmenting RNA
One of the more surprising recent developments is the discovery that Tn5 does not limit itself to double-stranded DNA. Researchers have shown that the enzyme can also tagment RNA/DNA hybrid duplexes, which are the structures formed during reverse transcription when an RNA molecule is being copied into its DNA complement. This finding opened the door to streamlined RNA sequencing workflows that skip one of the most tedious steps in traditional RNA-seq: synthesizing a second DNA strand before library preparation.
Two methods independently exploited this capability. SHERRY (Sequencing HEteRo RNA-DNA-hYbrid) showed that the Tn5 transposome can bind RNA/DNA hybrids and insert adapters directly onto the RNA strand after reverse transcription, making the method scalable and versatile.10PubMed Central. RNA sequencing by direct tagmentation of RNA/DNA hybrids A closely related approach called TRACE-seq similarly demonstrated direct tagmentation of RNA/DNA hybrids and was benchmarked against traditional RNA-seq for gene detection, gene body coverage, and the ability to measure differential expression.11eLife. Transposase-assisted tagmentation of RNA/DNA hybrid duplexes Protocols based on these methods have since been adapted for low-input RNA samples, making them practical for settings where material is scarce.12PubMed Central. Protocol for RNA-seq library preparation from low-volume total RNA by RNA/cDNA hybrid tagmentation
The ability to tagment RNA/DNA hybrids has also proven useful in viral diagnostics. Researchers applied the TRACE-seq method to prepare sequencing libraries from human norovirus samples and were able to recover nearly the entire viral genome, over seven kilobases, from all eleven clinical samples tested.13PubMed Central. Library Preparation Based on Transposase Assisted RNA/DNA Hybrid Co-Tagmentation for Next-Generation Sequencing of Human Noroviruses For RNA viruses, where obtaining full-length genomes is important for tracking mutations and evolution, this kind of streamlined approach is especially valuable.
Bead-Linked Tagmentation
A persistent frustration with standard tagmentation in solution is that the amount of enzyme must be carefully matched to the amount of input DNA. Too much enzyme for the DNA present means tiny fragments; too little means poor library yields. Quantifying your DNA accurately before starting can be a bottleneck in itself, especially in clinical or high-throughput settings where you want the workflow to be as hands-off as possible.
Bead-linked transposomes address this by conjugating a known quantity of Tn5 complexes directly to the surface of beads. When DNA is added, only a fixed amount binds to the bead-immobilized enzyme, regardless of how much DNA is in the tube. Library yield saturates at a certain input amount, around 100 nanograms in one commercial implementation, and below that threshold the method self-normalizes. The result is a sequencing-ready library from a wide range of DNA types and amounts, largely eliminating the need for careful DNA quantification upfront.14PubMed Central. Bead-linked transposomes enable a normalization-free workflow for NGS library preparation The technology even supports direct input of blood and saliva through an integrated extraction protocol, which is a significant simplification for clinical laboratories.
Beyond normalization, immobilizing transposomes on beads has a second advantage: it preserves long-range information. In solution-based tagmentation, once the enzyme cuts the DNA, the fragments drift apart and the information about which fragments were originally neighbors in the genome is lost. When tagmentation happens on a bead surface, the fragments stay physically tethered nearby, and this spatial relationship can be read out in the sequencing data. A method called tagmentation on microbeads (TOM) was specifically designed to exploit this, recovering long-range DNA sequence information that standard short-read sequencing normally destroys.15ACS Applied Materials & Interfaces. Tagmentation on Microbeads: Restore Long-Range DNA Sequence Information Using Next Generation Sequencing with Library Prepared by Surface-Immobilized Transposomes
Spatial Epigenomics and CUT&Tag
Tagmentation has moved beyond library preparation into more specialized roles in which the enzyme acts as a molecular probe. In a technique called CUT&Tag (Cleavage Under Targets and Tagmentation), the transposase is tethered to an antibody that recognizes a specific histone modification or other chromatin-associated protein. The enzyme sits idle until magnesium ions are added to activate it, at which point it inserts adapters only at the genomic sites where the target protein sits. This approach maps protein-DNA interactions with low cell input requirements and minimal background noise.
Taking this a step further, researchers have combined CUT&Tag chemistry with spatial barcoding to create tissue-level maps of histone modifications. In a method called hsrChST-seq, a tissue section on a glass slide is lightly fixed, treated with antibodies against a target histone mark, and then incubated with a protein A-Tn5 transposome fusion. When the transposome is activated by magnesium, it inserts barcoded adapters at the antibody binding sites. The spatial coordinates of each barcode are known in advance, so the sequencing data can be overlaid onto the tissue architecture, creating a map of where specific epigenetic marks exist in a tissue at high spatial resolution.16bioRxiv. Spatial Epigenome Sequencing at Tissue Scale and Cellular Level This represents a genuinely new capability: until recently, histone modification profiling required pooling thousands of cells, obliterating any spatial context.
Cost and Open-Source Access
Commercial Tn5 preparations have historically been expensive, and the proprietary formulations behind major platforms were undisclosed, which limited both cost-effective large-scale projects and the development of new applications.3PubMed Central. Tn5 transposase and tagmentation procedures for massively scaled sequencing projects Over the past several years, multiple groups have published protocols for purifying hyperactive Tn5 in standard molecular biology laboratories. These protocols yield multi-milligram quantities of functional enzyme per liter of bacterial culture, enough for thousands of tagmentation reactions, and the enzyme retains activity in storage for over a year.17PubMed Central. OpenTn5: Open-Source Resource for Robust and Scalable Tn5 Transposase Purification and Characterization
The availability of open-source Tn5 protocols has been quietly transformative for the field. Labs in resource-limited settings can now perform ATAC-seq, CUT&Tag, and standard library preparation at a fraction of the cost of commercial kits. It has also enabled rapid customization: researchers can load their own adapter sequences, experiment with novel barcoding strategies for single-cell work, and tailor the enzyme-to-DNA ratio for unusual sample types without being locked into a vendor’s recommended workflow. The combination of a well-characterized enzyme, published purification protocols, and low material costs means tagmentation has become one of the most accessible and adaptable tools in modern genomics.
Beyond Tn5
Tn5 dominates the tagmentation landscape, but it is not the only transposase that can be repurposed for sequencing applications. The MuA transposase, derived from bacteriophage Mu, has been explored as an alternative. MuA shows a remarkable tolerance for modifications to its transposon DNA sequence. Researchers have engineered MuA transposomes carrying randomized nucleotides in twelve positions of the transposon, creating an enormous diversity of molecular barcodes, theoretically around 10^17 unique variants, that are introduced into target DNA upon insertion.18Journal of Molecular Biology. MuA-based Molecular Indexing for Rare Mutation Detection by Next-Generation Sequencing This capacity for molecular indexing makes MuA attractive for applications such as rare mutation detection, where distinguishing true low-frequency variants from sequencing errors requires a unique barcode on every original DNA molecule.
The MuA system shares the GC insertion preference seen with Tn5, however, so it does not solve the bias problem on its own.6PubMed Central. Insertion site preference of Mu, Tn5, and Tn7 transposons Other transposon systems, such as Tn7, show far less sequence bias and more uniform insertion spacing, but they have not yet been as thoroughly adapted for high-throughput sequencing workflows. As the field continues to push toward single-cell and spatial applications that demand ever more uniform and efficient tagmentation, the search for better-behaved transposases remains an active line of engineering work.