An operational taxonomic unit, or OTU, is a label researchers use to group microorganisms by genetic similarity when they cannot easily assign each one to a named species. Because most microbes are impossible to grow in a lab and study under a microscope, scientists instead extract DNA from a sample, sequence a specific marker gene, and sort the resulting sequences into clusters. Each cluster is an OTU, and it serves as a practical stand-in for “a type of microbe.” The concept has become central to modern microbiology, shaping how we measure diversity in environments from ocean sediments to the human gut, even as newer methods begin to refine or replace it.
Why Microbes Need a Different Kind of Classification
For plants and animals, species are defined largely by whether organisms can breed with each other and produce fertile offspring. Microbes do not cooperate with that framework. Bacteria reproduce by dividing, they swap genetic material horizontally with unrelated neighbors, and the vast majority have never been cultured in a laboratory. That means there is no universally agreed-upon definition of a bacterial “species” in the way most people understand the word. Even closely related bacterial isolates can have large differences in genome content because of ongoing horizontal gene transfer and gene loss.
Early bacterial classification relied on observable traits like cell shape, staining behavior, and growth requirements. That approach worked for the handful of microbes you could culture, but it missed the enormous hidden majority. The shift toward DNA-based methods in the late twentieth century opened a window into that hidden diversity but introduced a new problem: if you have millions of DNA sequences from a soil sample, how do you decide which ones belong to “the same kind” of organism? OTUs were invented to answer that question in a standardized, reproducible way.
How OTUs Are Built
The process starts with a marker gene, almost always the gene encoding 16S ribosomal RNA in bacteria and archaea. This gene is useful because every bacterium has it, it contains regions that are highly conserved (good for designing universal primers that grab it from any species) alongside regions that vary enough to tell species apart. Researchers extract total DNA from a sample, use PCR to amplify the 16S gene, and then sequence the amplified fragments.
The sequencing run produces thousands to millions of short reads. Clustering software then groups those reads by similarity. Two sequences that are, say, at least 97% identical to each other get placed in the same bin. That bin is one OTU. The 97% similarity threshold was chosen with the intended purpose that the resulting OTUs could serve as a rough proxy for bacterial species.1PubMed Central. Reconciliation between operational taxonomic units and species boundaries It is not a law of nature but a pragmatic cutoff that, on average, separates most recognized species from one another when applied to the 16S gene.
Two broad strategies exist for the clustering step. In “de novo” clustering, sequences are compared only to each other, with no external reference. In “closed-reference” clustering, sequences are matched against a curated database of known 16S sequences, and anything that does not hit the database is thrown out. A hybrid called “open-reference” clustering tries the database first and then clusters leftover reads de novo, combining the speed of reference matching with the ability to capture novel organisms.2PubMed Central. Subsampled open-reference clustering creates consistent, comprehensive OTU definitions and scales to billions of sequences
What the 97% Threshold Actually Means
If two 16S sequences share 97% or more of their nucleotides, they land in the same OTU. The logic is that most pairs of sequences within a recognized bacterial species fall above that line, while most pairs drawn from different species fall below it. In practice, the relationship is messier. Some well-defined species have 16S sequences that are more than 99% identical to their nearest relative, meaning a 97% threshold lumps them together. Other species contain enough internal variation that a single species gets split into multiple OTUs.
Researchers have experimented with stricter cutoffs, such as 99%, and with looser ones. One comparison found that the choice between 97% and 99% identity thresholds had a surprisingly small effect on the broad ecological patterns detected, especially compared to the effect of choosing different analysis pipelines altogether.3PubMed Central. Ranking the biases: The choice of OTUs vs. ASVs in 16S rRNA amplicon data analysis has stronger effects on diversity measures than rarefaction and OTU identity threshold That does not mean the threshold is irrelevant, but it does mean that how you process the data matters at least as much as where you draw the similarity line.
What Researchers Do with OTU Tables
Once sequences are clustered, the result is an OTU table: a spreadsheet-like matrix listing every OTU found in each sample and how many sequence reads belong to it. That table becomes the raw material for measuring microbial diversity. Researchers typically look at two kinds. Alpha diversity asks how many different types live within a single sample and how evenly they are distributed. Beta diversity asks how different two samples are from each other. Both are estimated using the OTU table and standard ecological indexes.4PLOS ONE. Bifidobacteria Abundance-Featured Gut Microbiota Compositional Change in Patients with Behcet’s Disease
These diversity measures have driven a huge body of research. Studies of the human gut microbiome, for instance, use OTU-based profiling to compare the microbial communities of healthy people with those of patients who have inflammatory bowel disease, diabetes, or other conditions. In oral health, OTU profiling has identified distinct community types in people with gum disease, with different clusters of bacteria associated with different extents of periodontal damage.5PLOS ONE. Microbiome Profiles in Periodontitis in Relation to Host and Disease Characteristics Without a standardized unit like the OTU, comparing results across studies would be far harder.
Where OTUs Go Wrong
OTUs are a practical tool, and practical tools have practical problems. The biggest is that the clustering process can create phantom diversity that does not exist in the real sample. Errors introduced during PCR amplification and sequencing generate slightly altered copies of real sequences. Those erroneous reads may fall just outside the similarity threshold and get their own OTU, inflating the apparent richness of the community.
Chimeric sequences are an especially troublesome artifact. During PCR, an incomplete copy of one gene can fuse with a fragment from a different organism, producing a hybrid that looks like a novel species. In one controlled experiment using a mock community of known organisms, about 8% of the raw reads were chimeric. After quality filtering and chimera-detection software, the rate dropped to around 1%, but the remaining undetected chimeras were largely responsible for spurious OTUs that did not correspond to any real organism in the sample.6PLoS ONE. Reducing the Effects of PCR Amplification and Sequencing Artifacts on 16S rRNA-Based Studies The number of these ghost OTUs increased with sequencing depth, meaning that the more thoroughly you sequenced, the more fake diversity you generated.
PCR can also distort relative abundances. Some organisms’ 16S genes amplify more efficiently than others, so what looks like a dominant species in your data might just be a species whose DNA copies well in a test tube. That said, diversity metrics that account for both richness and relative abundance tend to be less affected by spurious OTUs than metrics based on richness alone.7PLoS ONE. PCR Biases Distort Bacterial and Archaeal Community Structure in Pyrosequencing Datasets
Stability and Reproducibility
A subtler problem is that OTU definitions can shift depending on the dataset. When you add new sequences to a de novo clustering run, the boundaries of existing OTUs can change because the algorithm is comparing everything against everything else. An OTU labeled “OTU_42” in one study is not guaranteed to correspond to the same group of organisms labeled “OTU_42” in another study, even if both used the same threshold and software. Closed-reference clustering is the only method that produces completely stable OTUs, because it anchors every cluster to a fixed reference database, but it discards any sequence that does not match the database, potentially missing novel organisms entirely.8PubMed Central. Stability of operational taxonomic units: an important but neglected property for analyzing microbial diversity
This instability has real consequences for cross-study comparisons. If two labs analyze the same environment but use different clustering strategies, reference databases, or even different random-number seeds in their software, their OTU lists can look different even though the underlying biology is identical. The field has spent considerable effort standardizing pipelines and database versions to reduce this noise, but it has never fully gone away.
The Rise of Amplicon Sequence Variants
Frustration with the limitations of OTU clustering fueled the development of a finer-grained alternative called amplicon sequence variants, or ASVs. Instead of grouping sequences into bins defined by a similarity threshold, ASV methods use statistical models to identify and correct sequencing errors, then report every unique corrected sequence individually. The result is resolution down to single-nucleotide differences.9The ISME Journal. Exact sequence variants should replace operational taxonomic units in marker-gene data analysis
The advantages are real. ASVs are consistent labels: a given ASV sequence means the same thing regardless of which dataset it appears in, because it is defined by its own nucleotide string rather than by its relationship to whatever other sequences happened to be in the run. That makes ASVs inherently more reproducible and easier to compare across studies. Proponents have argued that these improvements in reusability and comprehensiveness are significant enough that ASVs should replace OTUs as the default unit of marker-gene analysis.
That said, the practical differences are not always dramatic. One large field study found that broadscale ecological patterns were robust regardless of whether the analysis used ASVs or 97% OTUs.10PubMed Central. Broadscale Ecological Patterns Are Robust to Use of Exact Sequence Variants versus Operational Taxonomic Units The big-picture community differences between samples remained the same. A comparison of the two approaches across multiple datasets likewise found that the choice of pipeline (DADA2 for ASVs versus Mothur for OTUs) had a stronger effect on diversity measurements than the choice of OTU threshold or even whether rarefaction was applied.3PubMed Central. Ranking the biases: The choice of OTUs vs. ASVs in 16S rRNA amplicon data analysis has stronger effects on diversity measures than rarefaction and OTU identity threshold In other words, ASVs and OTUs often agree on the ecological story; the differences show up more at fine taxonomic scales and in cross-study comparisons. Many researchers now default to ASVs for new work, but the enormous existing literature built on OTUs remains valid and widely referenced.
Assigning Names to OTUs
An OTU by itself is just a numbered cluster. To give it a biological identity, researchers compare its representative sequence against a reference database of annotated 16S sequences. Several major databases exist, including Greengenes, SILVA, the Ribosomal Database Project (RDP), and NCBI’s nucleotide collection. Which database you use can change the names that come back. In one evaluation using a mock community of 60 known strains, the EzBioCloud database identified over 40 of the 44 expected genera across all samples, while Greengenes found only 30, and SILVA found a sufficient number but produced the highest rate of false positives, with around 20% of predicted genera being incorrect.11PubMed Central. Evaluation of 16S Databases for Taxonomic Assignments Using a Mock Community
Accuracy also drops sharply at finer taxonomic levels. Most databases do a reasonable job assigning OTUs to the right phylum or family, but species-level identification using a short fragment of the 16S gene is often unreliable. The 16S gene simply does not vary enough between some closely related species to tell them apart. Efforts to merge and curate sequences from multiple databases into unified collections have improved things somewhat, but species-level resolution remains a known weak spot of the entire marker-gene approach.12PubMed Central. Construction & assessment of a unified curated reference database for improving the taxonomic classification of bacteria using 16S rRNA sequence data
OTUs Beyond Bacteria
Although 16S-based OTUs are most closely associated with bacterial and archaeal surveys, the same logic applies to other groups of microorganisms. Fungal ecologists use the ITS (internal transcribed spacer) region of ribosomal DNA as their primary marker gene, clustering ITS sequences into fungal OTUs at various similarity thresholds. The choice of marker gene matters here too. A study of tidal estuary sediments found that an ITS2 primer set captured high diversity among common fungal groups but missed many early-diverging fungal lineages, while a 28S ribosomal primer set excelled at detecting those overlooked groups, with over half of the fungal OTUs it identified belonging to lineages the ITS primers had missed.13BioMed Central / Springer Nature (Environ Microbiome). Beyond dikarya: 28S metabarcoding uncovers cryptic fungal lineages across a tidal estuary The broader lesson is that no single marker gene captures everything, and the picture you get depends heavily on which gene you choose to amplify.
For eukaryotic microorganisms like protists and microalgae, the 18S ribosomal RNA gene fills a similar role to 16S in bacteria, though reference databases for eukaryotic microbes are generally less complete. Environmental DNA (eDNA) surveys of water and soil routinely produce OTU tables that span bacteria, fungi, and protists simultaneously, each identified through different marker genes processed through the same clustering logic.
When Shotgun Metagenomics Offers More
All marker-gene approaches, whether they produce OTUs or ASVs, share a fundamental limitation: they sequence only one gene. Shotgun metagenomics takes a different route by shredding all the DNA in a sample into fragments and sequencing everything, not just one marker. This captures far more taxonomic information and also reveals functional genes, giving insight into what the community can do rather than just who is present.
Comparisons between the two approaches consistently show that 16S-based methods detect only a subset of what shotgun sequencing reveals. In a study of chicken gut microbiota, 16S sequencing captured only part of the community that shotgun sequencing identified, and shotgun data had more power to detect less abundant taxa.14PubMed Central. Comparison between 16S rRNA and shotgun sequencing data for the taxonomic characterization of the gut microbiota A study of seagull gut microbiota found that the two methods agreed reasonably well at broad taxonomic levels like phylum and family, but diverged increasingly at finer levels. At the family level, shotgun metagenomics detected 118 taxa that 16S sequencing missed entirely.15PubMed Central. Comparative analysis of shotgun metagenomics and 16S rDNA sequencing of gut microbiota in migratory seagulls
So why does anyone still use 16S and OTUs? Cost and simplicity. Shotgun metagenomics requires far more sequencing data per sample, more computational resources for analysis, and more complex bioinformatic pipelines. For studies comparing dozens or hundreds of samples at a broad community level, 16S amplicon sequencing with OTU or ASV analysis remains a practical and well-validated choice.
OTUs in Environmental Monitoring
OTU-based community profiling has moved well beyond academic ecology labs. Environmental regulators and industry groups are exploring DNA-based monitoring as a complement to traditional methods that rely on identifying organisms under a microscope. In aquatic ecosystems, sequencing microbial communities and analyzing OTU tables can reveal shifts in biodiversity driven by pollution, habitat degradation, or climate change.16PLoS ONE. Phylogenetic and Functional Metagenomic Profiling for Assessing Microbial Biodiversity in Environmental Monitoring
Offshore oil drilling provides a concrete example. In a study of marine sediments around drilling sites on the Norwegian continental shelf, DNA metabarcoding of eukaryotic communities identified changes in sediment biodiversity near oil platforms that agreed with results from traditional morphology-based surveys. The DNA approach also flagged potential indicator taxa that responded to pollutants associated with drilling fluids, offering a more detailed picture of ecological disturbance than microscope counts alone could provide.17PubMed. High-throughput metabarcoding of eukaryotic diversity for environmental monitoring of offshore oil-drilling activities As sequencing costs continue to fall, OTU-based and ASV-based methods are likely to become routine tools in regulatory environmental assessment, supplementing rather than replacing visual identification of organisms.
Macroecological Patterns in Microbial Communities
One of the more striking findings to emerge from decades of OTU-based research is that microbial communities follow some of the same broad ecological rules that govern plant and animal communities. Species-abundance distributions, the way a few types dominate while many are rare, look remarkably similar whether you are counting tree species in a rainforest or OTUs in a gram of soil. Recent work has identified quantitative macroecological laws governing how microbial species fluctuate in abundance across communities and over time, and shown that a relatively simple mathematical model based on environmental randomness can predict those patterns.18Nature Communications. Macroecological laws describe variation and diversity in microbial communities These findings matter because they suggest that microbial ecosystems, despite their astronomical diversity and rapid generation times, are governed by the same kinds of forces that shape the rest of the living world. OTU tables, for all their imperfections, have been the primary data source making that kind of analysis possible.