Minimap2 is a fast, general-purpose software tool that aligns DNA and RNA sequences against a reference genome or database. Developed by Heng Li and first described in a 2018 paper in Bioinformatics, it handles an unusually wide range of input types, from short accurate reads to long, error-prone nanopore sequences to entire assembled chromosomes. That flexibility, combined with its speed, has made it one of the most widely used alignment tools in genomics and a near-default starting point for long-read sequencing analysis.
The Problem Minimap2 Solves
Sequencing machines do not read an entire genome in one piece. They produce millions of fragments, called reads, that range from a hundred bases to tens of thousands of bases long. To make sense of those fragments, researchers need to figure out where each one belongs in a known reference genome. That matching process is called alignment, and it is one of the most computationally demanding steps in any genomics workflow. A human genome has roughly three billion base pairs; aligning millions of reads against it requires software that can search efficiently and tolerate the inevitable mismatches, insertions, and deletions found in real sequencing data.
Before minimap2, different sequencing technologies often required different alignment tools. Short-read aligners were optimized for reads of a few hundred bases with low error rates. Long-read aligners had to tolerate much higher error rates but were not designed for short reads. Minimap2 collapsed that divide. It can map accurate short reads of 100 bases or longer, long genomic reads of a kilobase or more with error rates around 15%, full-length noisy direct RNA or cDNA reads, and even assembled contigs or closely related chromosomes hundreds of megabases in length.1PubMed Central. Minimap2: pairwise alignment for nucleotide sequences That range of compatibility is what earned it the “versatile” label.
How the Alignment Process Works
At a high level, minimap2 works in stages. First, it builds an index of the reference genome by extracting short sequence fragments called minimizers, which serve as compact signatures that allow rapid lookup. When a query read comes in, minimap2 finds matching minimizers between the read and the reference, identifies clusters of matches that suggest a candidate mapping location, and then chains those matches together to determine the best alignment. A final base-level alignment step refines the result, accounting for mismatches and gaps.
Two design choices stand out. Minimap2 supports split-read alignment, which means it can handle reads that span structural breakpoints or exon-intron boundaries by splitting a single read into two or more aligned segments. It also uses a concave gap cost model, which penalizes long insertions and deletions more realistically than a simple linear penalty would. In biological sequences, a single large deletion is more likely than many small ones adding up to the same length, and the concave model reflects that.1PubMed Central. Minimap2: pairwise alignment for nucleotide sequences These features let minimap2 produce cleaner alignments for the kinds of complex variation that long reads frequently reveal.
Presets for Different Data Types
One reason minimap2 caught on so quickly is that it ships with built-in presets tuned for specific sequencing scenarios. Rather than asking the user to manually adjust dozens of parameters, it offers command-line flags that configure everything at once. The most commonly used presets include options for Oxford Nanopore (ONT) genomic reads, PacBio HiFi reads, short Illumina-style reads, spliced RNA alignment, and whole-genome assembly comparison. Choosing the right preset adjusts things like the minimizer window size, the gap penalties, and whether the tool looks for splice sites.
For RNA sequencing, the spliced alignment preset tells minimap2 to expect large gaps corresponding to introns and to look for canonical splice-site signals (the GT-AG dinucleotides at intron boundaries). This makes it useful for mapping long-read RNA data directly to a reference genome. In nanopore direct RNA sequencing workflows, for instance, reads are routinely aligned with minimap2 as a first step before downstream tools analyze RNA modifications or expression levels.2Nature Communications. RNA modifications detection by comparative Nanopore direct RNA sequencing That said, specialized RNA aligners such as uLTRA have been shown to outperform minimap2 on certain edge cases, particularly for small exons, where the general-purpose tool can miss the correct splice boundaries.3Bioinformatics. Accurate spliced alignment of long RNA sequencing reads
Speed and Computational Design
Minimap2 was built with speed in mind. It processes query sequences in batches, with each sequence in a batch handled independently in parallel. This batch-based design keeps memory usage manageable even when aligning against large genomes.4PubMed Central. Accelerating minimap2 for whole-genome alignment On a typical multi-core server, minimap2 can align a full set of human whole-genome nanopore reads in hours rather than days.
The batch approach does have a drawback: if a batch contains fewer query sequences than available CPU threads, some threads sit idle. Researchers have explored ways to address this inefficiency, including extracting more parallelism from the chaining step itself so that individual long sequences can be processed across multiple threads simultaneously.5PubMed Central. Accelerating Minimap2 for Accurate Long Read Alignment on GPUs Hardware acceleration has also been pursued. One FPGA-based implementation, minimap2-fpga, achieved up to 79% faster mapping for ONT datasets and up to 53% faster for PacBio datasets compared with the standard software version, with near-identical accuracy.6Nature Publishing Group. Efficient end-to-end long-read sequence mapping using minimap2-fpga integrated with hardware accelerated chaining These efforts reflect how central minimap2 has become: rather than replacing it, teams are investing in making it faster on new hardware.
Where Minimap2 Fits in Analysis Pipelines
Alignment is rarely the end goal. It is a stepping stone to answering biological questions, and minimap2 feeds into a wide variety of downstream analyses. One of the most prominent is structural variant (SV) calling, where researchers look for large-scale rearrangements in the genome such as deletions, duplications, inversions, and translocations. Long reads are especially good at detecting these because they can span entire variant breakpoints in a single read, and minimap2’s split-read alignment is well suited to flagging them.
A recent benchmarking study of long-read SV detection pipelines found that minimap2 paired with the SV caller Sniffles forms a strong baseline due to their combined speed and solid performance. In that comparison, the CuteSV caller achieved the highest average F1-score of about 83% and recall of roughly 79%, while Sniffles reached the highest average precision at about 94%.7PubMed Central. Benchmarking long-read aligners and SV callers for structural variation detection in Oxford nanopore sequencing data The alignment step was consistent across these caller comparisons, underscoring minimap2’s role as reliable infrastructure rather than a bottleneck.
Beyond structural variants, minimap2 alignments feed into base modification detection, gene expression quantification from RNA-seq, de novo genome assembly polishing, and metagenomic classification. In nanopore direct RNA sequencing, for example, minimap2 aligns reads to a reference so that downstream tools can compare the electrical signal to expected patterns and identify chemical modifications like m6A. One systematic comparison of m6A detection tools noted that reads from samples containing the modification had slightly lower alignment identity, likely because m6A causes subtle shifts in the nanopore signal that base-callers interpret as errors.8Nature Communications. Systematic comparison of tools used for m6A mapping from nanopore direct RNA sequencing Minimap2 handles these noisy reads without special configuration, which is part of why it became the default aligner in so many nanopore workflows.
Known Limitations and How They Have Been Addressed
No alignment tool is perfect, and minimap2 has well-documented weak spots. The most significant involves highly repetitive regions of the genome. Minimap2 relies on minimizers as seeds for alignment. During mapping, it originally selected only low-occurrence minimizers, filtering out those that appear too frequently in the reference. The cutoff was typically a few hundred occurrences for a human genome. If a read fell entirely within a highly repetitive region and contained few or no low-occurrence minimizers, the tool could fail to chain enough anchor points together, resulting in a missed or incorrect alignment.
Version 2.22 introduced a targeted fix for this. When two adjacent low-occurrence minimizers are far apart (more than 500 bases by default), minimap2 now selects additional minimizers of the lowest available occurrence from the intervening region. This fills in gaps in the anchor chain at the cost of only a small increase in total alignment time.9Bioinformatics. New strategies to improve minimap2 alignment accuracy The improvement helped, but repetitive regions remain a challenge for the tool, especially in genomes with extensive segmental duplications or centromeric repeats.
Another limitation is that minimap2’s heuristic approach occasionally produces spurious alignments, particularly for reads that originate from regions not present in the reference. The original paper acknowledged this and introduced heuristics to reduce such artifacts, but they have not been entirely eliminated. For applications where alignment accuracy in difficult regions is critical, researchers sometimes turn to specialized tools or combine minimap2 with post-alignment filtering steps.
Winnowmap and Other Tools That Build on Minimap2
Minimap2’s open-source codebase and modular design have made it a foundation for derivative tools. The most notable is Winnowmap, which directly addresses the repetitive-region weakness. Winnowmap uses a weighted minimizer sampling strategy: it assigns lower weights to highly repetitive k-mers (those appearing above a frequency cutoff, by default 1,024 times in the reference) so that they are less likely to be selected as seeds. This avoids excessive false-positive matches while still maintaining the guarantee that nearby sequences will share at least one minimizer.10PubMed Central. Weighted minimizer sampling improves long read mapping
Winnowmap2 extended this further by introducing the concept of minimal confidently alignable substrings. Instead of attempting to align a full read in one pass, it breaks the read into sub-alignments that can each be placed with high confidence, then combines them. This makes the tool more tolerant of structural variation between the read and the reference and more sensitive to paralog-specific variants within repetitive sequences. The result is reduced allelic bias and more accurate downstream variant calls in regions where minimap2 tends to struggle.11Nature Methods. Long-read mapping to repetitive reference sequences using Winnowmap2 These tools are not replacements for minimap2 so much as complements: they share much of its architecture but are specifically tuned for genomes or genomic regions where repetitive content dominates.
Why Minimap2 Became the Default
Software tools in bioinformatics come and go quickly. A new aligner appears every year or two, each claiming improvements in speed, accuracy, or both. Minimap2 has endured because it hit a useful sweet spot. It was fast enough that researchers did not need to wait overnight for results, accurate enough for the vast majority of applications, and flexible enough that a single tool could handle reads from Illumina, PacBio, and Oxford Nanopore platforms. That meant fewer dependencies, simpler pipelines, and less time spent learning platform-specific software.
Its adoption was also boosted by timing. Minimap2 arrived just as long-read sequencing was transitioning from a niche technology to a mainstream one. The Oxford Nanopore MinION had become affordable enough for individual labs, and PacBio’s HiFi chemistry was producing reads with dramatically improved accuracy. Researchers needed an aligner that could keep up with both the volume and the diversity of these new data types. Minimap2 filled that gap before any competitor did, and network effects took over: pipelines were written around it, tutorials featured it, and benchmarking studies used it as the baseline comparison.
The community around the tool has been active as well. Heng Li has continued releasing updates that incorporate new heuristics and optimizations, and the codebase has been extended by other groups for GPU acceleration, FPGA integration, and domain-specific applications. Even large-scale genome assembly projects, such as the nanopore-based assembly of eleven human genomes using the Shasta toolkit, have relied on minimap2 or tools derived from its algorithms for read-to-reference and read-to-read overlap computations.12Nature Biotechnology. Nanopore sequencing and the Shasta toolkit enable efficient de novo assembly of eleven human genomes
Practical Tips for Getting Started
If you are new to minimap2, the learning curve is gentler than you might expect. It runs from the command line on Linux and macOS, takes a reference genome and a set of reads as input, and outputs alignments in the standard SAM or PAM format that nearly every downstream tool can read. Installation is straightforward through package managers like conda or by compiling from the source code on GitHub.
The most important decision you make is selecting the right preset. For Oxford Nanopore genomic reads, the map-ont preset is standard. For PacBio HiFi reads, map-hifi is the choice. For spliced RNA alignment, splice mode engages the intron-aware logic. Picking the wrong preset will not crash the tool, but it will degrade alignment quality because the internal parameters will be tuned for the wrong error profile and read-length distribution. If your data does not fit neatly into one of the built-in categories, the minimap2 manual documents every parameter you can adjust individually, though most users never need to go that far.
One practical consideration is memory. Indexing a human reference genome requires several gigabytes of RAM, and aligning a full dataset against it needs more. The batch-processing design keeps this manageable on most modern servers, but if you are working on a laptop or a shared computing cluster with tight memory limits, you may need to split your input or reduce the index size. For smaller genomes, such as bacterial or viral references, minimap2 runs comfortably on modest hardware.
When to Use Something Else
Minimap2 is a strong default, but it is not always the best tool for every job. For short-read alignment specifically, BWA-MEM2 and Bowtie2 remain popular and are optimized for the particular characteristics of Illumina data. If your reads are all short and accurate, these dedicated short-read aligners can offer slightly better sensitivity for certain variant types, particularly small insertions and deletions.
For RNA-seq data with complex splicing patterns, particularly when detecting small exons is important, specialized tools like uLTRA or STAR may provide higher accuracy than minimap2’s splice preset. The trade-off is usually speed or ease of use: minimap2 is simpler to set up and runs faster, while the specialist tools can squeeze out better results in difficult cases.
For genomes rich in segmental duplications or centromeric repeats, Winnowmap2 is the more appropriate choice, as described above. And for applications requiring base-level consensus accuracy, such as polishing draft genome assemblies, minimap2 provides the alignment but specialized polishing tools like Medaka or DeepVariant handle the actual error correction. Minimap2 is infrastructure. It gets your reads to the right place in the genome; what you do with them there depends on the question you are asking.