What Is Kallisto and How Does It Work?

Kallisto is a free, open-source bioinformatics program that quantifies how much each gene (or more precisely, each transcript) is expressed in an RNA sequencing experiment. Published in Nature Biotechnology in 2016, it introduced a concept called pseudoalignment that made it roughly a hundred times faster than the methods researchers had been using, while achieving similar accuracy. The speed gain is dramatic enough that a standard laptop can chew through 30 million paired-end RNA-seq reads in under ten minutes, a task that previously demanded powerful servers and hours of compute time.

The Problem Kallisto Was Built to Solve

RNA sequencing generates millions of short DNA fragments, called reads, that correspond to the RNA molecules active in a cell or tissue at a given moment. To figure out which genes are turned on and by how much, researchers need to trace each read back to the transcript it came from. The traditional way to do this is alignment: comparing every single base in every read against a reference genome or transcriptome and finding the best positional match. Tools built around full alignment are thorough, but they are computationally expensive. They require significant memory, powerful hardware, and patience. As sequencing became cheaper and datasets ballooned, that bottleneck became a real obstacle. A lab generating dozens of samples a week could easily spend more time waiting for software to finish than it spent actually collecting data.

Kallisto’s insight was that for transcript quantification, you do not actually need to know where every base in a read lands on the genome. You just need to know which transcript or set of transcripts a read is compatible with. That shift in thinking opened the door to a much faster approach.

How Pseudoalignment Works

Pseudoalignment skips the step of aligning individual bases to a reference. Instead, it produces a list of transcripts that are compatible with each read, without pinpointing exactly where along a transcript the read sits. The process relies on a data structure built from the reference transcriptome called a transcriptome de Bruijn graph, or T-DBG. This graph encodes all the short subsequences (called k-mers) found across every known transcript, along with information about which transcripts share which sequences.

When a read comes in, kallisto breaks it into k-mers and looks up each one in the graph. The goal is to find the set of transcripts that every k-mer in the read is consistent with. In practice, kallisto uses a shortcut to avoid checking every single k-mer: it examines k-mers at the beginning, middle, and end of the read first. If those all point to the same set of transcripts, no further checking is needed. Only when those initial lookups disagree does the software examine additional k-mers in a more thorough, iterative fashion. The result of this process is a transcript compatibility class, which is simply the group of transcripts a given read could plausibly have come from.

This is the key trade-off. Pseudoalignment sacrifices positional information (where exactly the read maps within a transcript) in exchange for enormous speed. For the specific task of counting how many reads belong to each transcript, positional detail turns out to be largely unnecessary.

From Compatibility Classes to Abundance Estimates

Knowing which transcripts a read could have come from is not the same as knowing which transcript it actually came from. Many reads are ambiguous: they match two or more transcripts equally well, usually because those transcripts share stretches of identical sequence. Kallisto resolves this ambiguity using an expectation-maximization (EM) algorithm, a statistical method that iteratively refines its best guess about how many reads belong to each transcript.

The algorithm starts by distributing ambiguous reads evenly among their candidate transcripts, then adjusts those assignments based on the overall pattern of evidence. Transcripts that have a lot of uniquely mapped reads pull more of the ambiguous reads toward them in each round of estimation, and the process repeats until the numbers stabilize. The final output is a set of transcript-level abundance estimates, typically reported as transcripts per million (TPM), along with estimated counts. Researchers have found that TPM values from kallisto and similar tools show high linearity across samples, meaning the numbers scale reliably when comparing expression levels between conditions.

Speed, Memory, and What They Mean in Practice

The headline performance claim from the original kallisto paper is that it runs about two orders of magnitude faster than traditional alignment-based quantification, and independent benchmarks support that general picture. In one evaluation using single-cell RNA-seq data, kallisto took roughly a quarter of the time that STAR, a widely used traditional aligner, needed to process the same dataset. Memory use was even more lopsided: kallisto used about 3.6 gigabytes of RAM, compared to roughly 28 gigabytes for STAR.

Those numbers matter beyond mere convenience. A tool that runs in a few gigabytes of RAM can operate on a personal laptop. A tool that demands 28 gigabytes or more typically requires a dedicated computing cluster or a cloud instance that costs money by the hour. For individual researchers, small labs, or groups working in resource-limited settings, this difference can determine whether an analysis is feasible at all. Kallisto’s light footprint also makes it practical to rerun analyses quickly when parameters change or new samples arrive, which encourages more exploratory and iterative science.

How Accurate Is It Compared to Traditional Aligners

Speed is useful only if the results hold up, and the evidence here is encouraging but nuanced. On idealized simulated data, kallisto performs among the most accurate tools available, grouped with Salmon, RSEM, and Cufflinks at the top of evaluations. On more realistic data that includes the noise and messiness of actual sequencing experiments, the gap between these top-tier tools and simpler approaches narrows: the sophisticated methods do not dramatically outperform more basic ones.

A systematic comparison of RNA-seq pipelines found that pseudoalignment-based tools like kallisto have good precision in gene expression estimation, but their accuracy can be somewhat lower than classic alignment methods. The study suggested that pseudoaligners are a strong choice as exploration tools, where speed lets researchers quickly survey their data and identify patterns worth investigating further. When the stakes demand the highest possible accuracy on individual genes, a traditional aligner may still have an edge in some cases.

In practice, the accuracy differences tend to be small enough that kallisto has become a default first-pass tool in many labs. Researchers often run kallisto to get a fast overview of their data, then turn to heavier methods selectively if a particular finding needs extra scrutiny.

Single-Cell RNA-Seq and the kb-python Ecosystem

Kallisto’s speed advantage becomes even more relevant in single-cell RNA sequencing, where a single experiment can generate data from thousands or tens of thousands of individual cells. Each cell produces its own set of reads, tagged with molecular barcodes that identify which cell and which original RNA molecule each read came from. Processing this data with traditional aligners can be extremely time-consuming.

To handle single-cell workflows, the developers of kallisto created a companion tool called bustools and a Python wrapper called kb-python. Together, these tools form an integrated pipeline: kallisto handles the pseudoalignment, bustools processes the barcoded read data into a structured format (called BUS format, for Barcode, UMI, Set), and kb-python ties everything together so a user can go from raw sequencing files to a gene-by-cell expression matrix in a small number of commands. The workflow handles both cell-based assays (where RNA is captured from intact cells) and nucleus-based assays (where RNA is captured from isolated nuclei), and it can distinguish between newly made (nascent) RNA and fully processed (mature) RNA.

This ecosystem has made single-cell preprocessing accessible to researchers who do not have extensive computational resources. The entire pipeline is free and open-source, and because it inherits kallisto’s low memory requirements, it runs on standard hardware that most labs already own.

Metagenomics and Other Applications Beyond Gene Expression

The pseudoalignment concept turned out to be useful well beyond its original RNA-seq application. Researchers have adapted it for metagenomics, the study of all the microbial genomes present in an environmental or clinical sample. In metagenomics, the core task is similar: you have millions of sequencing reads, and you need to figure out which organisms they came from and in what proportions.

A study comparing metagenomics analysis pipelines found that a kallisto-based approach achieved the highest accuracy in estimating species-level abundances as sequencing depth increased. The best-performing pipeline used a multi-step strategy: first filtering out human reads, then detecting which species were present using marker genes, and finally quantifying abundances against a database of full genomes from just the detected species. This filtered approach outperformed established metagenomics tools in quantification accuracy, though it required more processing time than some faster but less accurate alternatives.

The underlying reason pseudoalignment transfers so well to metagenomics is that the mathematical problem is structurally the same. Whether you are asking “which transcript did this read come from?” or “which bacterial genome did this read come from?”, you are assigning reads to reference sequences and estimating relative abundances. The EM framework handles the ambiguity in both cases.

Long Reads and lr-kallisto

Kallisto was originally designed for short reads, the 75-to-300-base-pair fragments produced by the most common sequencing platforms. But sequencing technology has been moving toward long reads, which can span thousands of base pairs and sometimes capture entire transcripts in a single read. Long reads offer advantages for identifying transcript isoforms, but they also have higher error rates than short reads, which creates new challenges for any analysis tool.

The kallisto developers addressed this with lr-kallisto, an extension that adapts the pseudoalignment framework for long-read data. The core logic is the same, but several adjustments account for the differences in read characteristics. When a long read’s k-mers point to conflicting sets of transcripts (which happens more often with long reads because of their higher error rates), lr-kallisto takes the most frequently occurring transcript compatibility class rather than requiring a strict intersection. If at least one k-mer maps uniquely to a single transcript, lr-kallisto prioritizes the compatibility classes from those uniquely mapping k-mers.

The EM algorithm also received modifications for long reads. Uniquely mapping reads are handled differently during the estimation process: instead of being included from the start, they are added after the EM algorithm has already resolved the ambiguous reads, which improves the accuracy of the final abundance estimates. Testing showed that lr-kallisto is robust across a range of error rates, making it suitable for both newer, lower-error long-read platforms and older datasets with higher error rates. Length normalization, which adjusts for the fact that longer transcripts generate more reads, helped at low error rates but actually hurt performance when errors were high or unevenly distributed across reads.

Working with Downstream Analysis Tools

Kallisto’s output is designed to feed into established statistical frameworks for identifying differentially expressed genes. One natural partner is sleuth, a program built specifically to work with kallisto’s bootstrap estimates. Kallisto can generate bootstrap samples during quantification, which capture the uncertainty in each transcript’s abundance estimate. Sleuth uses these bootstraps to account for technical variability when testing whether a gene is expressed at different levels between experimental conditions.

Researchers also commonly pipe kallisto output into more general tools like DESeq2. In one validated workflow, transcript-level counts from kallisto are used as input for differential expression testing at the transcript level, and then the resulting statistical significance values are aggregated up to the gene level using a method called the Lancaster approach. This lets researchers benefit from transcript-resolution quantification while still getting gene-level results, which are often more interpretable for biological questions.

The practical upshot is that kallisto is not a standalone analysis tool. It handles one specific step, quantification, and does it quickly. Everything that happens after quantification, including normalization, statistical testing, and biological interpretation, relies on separate software. This modularity is a feature, not a limitation: it means researchers can swap in kallisto as the quantification engine while keeping whatever downstream pipeline they already trust.

Where Kallisto Falls Short

Pseudoalignment’s speed comes from discarding positional information, and that trade-off has consequences. If your question requires knowing exactly where reads map within a transcript or across the genome, kallisto cannot help. Studies of alternative splicing at base-pair resolution, structural variant detection, or any analysis that depends on read pileups at specific genomic coordinates will still need a traditional aligner.

Kallisto also depends entirely on the reference transcriptome you give it. If a transcript is missing from the reference, kallisto will never find it. Reads from unannotated transcripts either get incorrectly assigned to annotated ones that share some sequence, or they simply fail to pseudoalign. The lr-kallisto team has noted that reads failing to pseudoalign could, in principle, be assembled to discover unannotated transcripts, but that functionality has not been fully developed or benchmarked yet.

Genome complexity and gene family redundancy also create challenges. When many transcripts share long stretches of identical sequence, a large fraction of reads end up in ambiguous compatibility classes. The EM algorithm does its best to resolve these, but the estimates for highly similar transcripts inevitably carry more uncertainty. For organisms with well-annotated, relatively compact transcriptomes, this is rarely a major issue. For organisms with large gene families or poorly characterized genomes, it can degrade performance.

How Kallisto Compares to Salmon

Any discussion of kallisto eventually lands on Salmon, a tool developed independently around the same time that also uses a lightweight, alignment-free approach to transcript quantification. The two are frequently treated as near-interchangeable in practice, and benchmarks consistently show them clustered together in both speed and accuracy. Both use EM-based quantification, both bypass full alignment, and both produce TPM and count estimates that show high linearity across samples.

The differences are mostly architectural. Salmon uses a slightly different indexing strategy and offers features like sequence-specific bias correction and GC-content bias modeling out of the box. Kallisto’s ecosystem is more tightly integrated with single-cell tools through bustools and kb-python. In comparative evaluations, neither tool consistently outperforms the other across all datasets and conditions. The choice often comes down to which ecosystem a lab is already embedded in, or which tool’s documentation and community support better fits a researcher’s workflow. Researchers switching from one to the other typically see negligibly different results for the vast majority of genes.

The Index and What Goes Into It

Before kallisto can quantify anything, it needs to build an index from a reference transcriptome file. This is a one-time step for each organism and annotation version. The index encodes the k-mer composition of every transcript into the T-DBG structure that pseudoalignment queries. Building the index takes a few minutes for a typical mammalian transcriptome and produces a file of a few gigabytes.

The choice of reference matters more than many users realize. Using a transcriptome annotation that is incomplete or outdated can systematically bias results, because any transcript absent from the index is invisible to kallisto. For well-studied organisms like human and mouse, standard annotation sources like GENCODE or Ensembl provide comprehensive references. For less-studied organisms, researchers sometimes need to supplement annotations with their own transcript assemblies, which introduces its own uncertainties.

The k-mer length used in the index (31 bases by default) also affects the balance between sensitivity and specificity. A shorter k-mer increases the chance of finding matches but also increases ambiguity. A longer k-mer is more specific but may miss reads with sequencing errors near the k-mer boundaries. The default works well for typical short-read data, but users working with unusually short reads or high-error-rate data sometimes adjust it.

Who Uses Kallisto and Why It Caught On

Kallisto’s adoption was rapid after its 2016 publication, and its user base spans academic biology, clinical genomics, and pharmaceutical research. Several factors beyond raw speed drove uptake. The tool runs on Linux, macOS, and Windows, and its command-line interface is straightforward enough that a biologist with basic terminal skills can operate it. The low memory requirement meant that labs did not need to invest in new infrastructure. And because pseudoalignment is conceptually simple compared to the heuristics used by traditional aligners, many users felt more confident that they understood what the software was actually doing with their data.

The availability of kb-python extended kallisto’s reach to the rapidly growing single-cell community, where preprocessing had been a persistent pain point. Before tools like this existed, single-cell researchers often spent days wrestling with alignment and barcode-demultiplexing pipelines. Having a streamlined, well-documented solution that ran quickly on modest hardware removed a significant barrier to entry, especially for labs new to single-cell work.