How Are Microbes Classified? From Traits to Genetics

Microbes are classified through a layered combination of observable traits and genetic information, with DNA-based methods increasingly driving the system. For most of the history of microbiology, researchers sorted organisms by what they could see and measure: shape under a microscope, how cells responded to chemical stains, what they ate, what they produced. That approach still matters in clinical labs and field work, but the core framework for deciding which microbe is which now rests on comparisons of DNA and protein sequences. The shift has been dramatic enough that roughly half of all bacterial classifications have been revised since genome-scale methods took over.

What Phenotypic Classification Looks Like in Practice

Before DNA sequencing was practical, microbiologists relied on a toolkit of physical and biochemical observations. The most familiar example is the Gram stain, developed in the 1880s and still used every day in hospitals worldwide. It divides bacteria into two broad camps based on how their cell walls interact with crystal violet dye: Gram-positive bacteria retain the purple stain because of their thick cell wall layer, while Gram-negative bacteria lose it during a wash step and pick up a pink counterstain instead.1PubMed Central. Classification of two species of Gram-positive bacteria through hyperspectral microscopy coupled with machine learning That single test immediately narrows a doctor’s choice of antibiotic, because the two groups differ in drug susceptibility.

Beyond staining, classical microbiology built up a catalog of traits: cell shape (rod, sphere, spiral), whether an organism needs oxygen, what sugars it can ferment, whether it forms spores, how it moves, what enzymes it secretes. These features were compiled into identification keys and used to assign organisms to genera and species. The system worked reasonably well for common pathogens, but it had real blind spots. Many microbes look identical under a microscope yet behave very differently. Others change their appearance depending on growth conditions. And for organisms that refused to grow in the lab at all, phenotypic classification was simply impossible.

Chemical Fingerprints as Taxonomic Clues

Sitting between old-fashioned trait observation and modern genomics is chemotaxonomy, which classifies organisms by their chemical makeup rather than their genes. One widely used approach analyzes the fatty acids in a cell’s membrane. Different groups of microbes build their membranes from different combinations of fatty acids, and those profiles can serve as a kind of chemical fingerprint. Fatty acid profiling has proven useful for distinguishing closely related cyanobacteria, for instance, where visual identification is unreliable.2PubMed. Chemotaxonomy of heterocystous cyanobacteria using FAME profiling as species markers

These chemical signatures can also separate larger categories. Eukaryotic cells (organisms with nuclei, including fungi and protists) generally contain polyunsaturated fatty acids and steroids in their membranes, while prokaryotic cells (bacteria and archaea) typically do not.3PubMed Central. Rapid detection of taxonomically important fatty acid methyl ester and steroid biomarkers using in situ thermal hydrolysis/methylation mass spectrometry (THM-MS) A related technology, MALDI-TOF mass spectrometry, fires a laser at a microbial sample to generate a characteristic protein spectrum. That spectrum acts like a barcode: match it against a database and you can identify the organism in minutes. Recent work has even explored using MALDI-TOF to detect drug resistance patterns in tuberculosis strains based on differences in their spectra.4Crescent Journal of Medical and Biological Sciences. Evaluation of Antibiotic Resistance in Mycobacterium tuberculosis Clinical Strains by Culture-Based Antibiogram, Target Gene Sequencing, and Matrix-Assisted Laser Desorption Ionization Time of Flight Mass Spectrometry (MALDI-TOF MS) Assay

The 16S Gene That Rewrote the Tree of Life

The biggest single leap in microbial classification came from a molecule that every bacterium and archaeon carries: the 16S ribosomal RNA gene. This gene encodes part of the cellular machinery that builds proteins, and because protein synthesis is so fundamental, the gene has been conserved across billions of years of evolution. Carl Woese exploited this conservation in the 1970s, using ribosomal RNA sequences to propose three major domains of life: Bacteria, Archaea, and Eukarya.5PubMed Central. Carl Woese’s vision of cellular evolution and the domains of life Before Woese’s work, archaea were lumped in with bacteria. Ribosomal sequences showed they were as different from bacteria as we are.

The 16S gene works for classification because it contains nine stretches of highly variable sequence (labeled V1 through V9) sandwiched between more stable, conserved regions.6PubMed Central. A detailed analysis of 16S ribosomal RNA gene segments for the diagnosis of pathogenic bacteria The conserved parts let researchers design universal primers that grab the gene from nearly any bacterium; the variable parts let them tell organisms apart. Not all variable regions perform equally, though. The V4 region, for example, identified all 17 bacteria tested at the family level in one systematic comparison, but only about 82% at the genus level. The V9 region fared worse, finding fewer than 60% of bacteria at either level.7PLOS ONE. Development of an Analysis Pipeline Characterizing Multiple Hypervariable Regions of 16S rRNA Using Mock Samples Which region to sequence depends on the question being asked and the organisms expected in a sample.

How Fungi Get Their Own Barcode

The 16S gene is a bacterial and archaeal tool. Fungi need a different marker, and the one that has emerged as the standard is the internal transcribed spacer (ITS) region of ribosomal DNA. A large-scale comparison found that ITS has the highest probability of successfully identifying fungi across the broadest range of species, with the clearest gap between variation within a species and variation between species.8PubMed Central. Nuclear ribosomal internal transcribed spacer (ITS) region as a universal DNA barcode marker for Fungi It is not perfect everywhere: for some early-diverging fungal lineages and for ascomycete yeasts, a different ribosomal region (the large subunit) does better. But ITS remains the default starting point for fungal identification.

ITS sequencing has also been used to untangle the relationships among anaerobic gut fungi, a group that is difficult to classify by appearance alone. Phylogenetic analysis of ITS sequences clearly separated genera that had previously been defined by whether their cells had one or many reproductive structures.9PubMed. Identification and characterization of anaerobic gut fungi using molecular methodologies based on ribosomal ITS1 and 185 rRNA The molecular data confirmed what morphology had suggested, but with far more precision and reproducibility.

Fungal taxonomy carried an unusual historical burden: many fungi were given two separate scientific names, one for their sexual form and one for their asexual form, because the two stages often look completely different and were discovered independently. DNA sequencing made it clear these were the same organism, and mycologists eventually moved toward a “one fungus, one name” policy, abandoning the confusing dual system that had persisted for over a century.10PubMed Central. One Fungus = One Name: DNA and fungal nomenclature twenty years after PCR11PubMed Central. One fungus, one name promotes progressive plant pathology

When a Single Gene Is Not Enough

A single marker gene can only tell you so much. Two bacteria might share nearly identical 16S sequences yet differ by thousands of genes elsewhere in their genomes. This is where whole-genome comparisons enter the picture. For decades, the gold standard for deciding whether two bacterial strains belong to the same species was DNA-DNA hybridization, a laboratory technique that physically measures how well the total DNA from two organisms sticks together. It served the field for about 50 years, but it was slow, hard to reproduce, and impossible to apply to organisms that could not be grown in culture.

The replacement is average nucleotide identity (ANI), which compares the actual letter-by-letter similarity of two genomes using computer algorithms. ANI mirrors the old hybridization results closely, with the boundary for a prokaryotic species landing at roughly 95–96% identity.12PubMed Central. Shifting the genomic gold standard for the prokaryotic species definition The advantage is enormous: ANI values are digital, reproducible, and can be calculated from genomes retrieved from databases without ever touching a petri dish.

Taking this further, the Genome Taxonomy Database (GTDB) project used a phylogeny built from 120 concatenated protein-coding genes to construct a standardized bacterial taxonomy. The results were eye-opening: 58% of the roughly 95,000 genomes in the database had their existing classification changed.13Nature Biotechnology. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life Groups that had been lumped together turned out to be distantly related; others that had been split apart were actually close cousins. The accompanying software toolkit, GTDB-Tk, lets researchers classify new genomes against this reference tree, though the growing size of the bacterial tree has pushed memory requirements into the hundreds of gigabytes.14PubMed Central. GTDB-Tk v2: memory friendly classification with the genome taxonomy database

Classifying the Unculturable Majority

Most microbes on Earth have never been grown in a lab. Estimates vary, but the fraction that resists standard culturing is large, perhaps the vast majority of species in soil and ocean environments. For a long time, these organisms existed as dark matter in the microbial world: known only from environmental DNA sequences, impossible to classify by traditional means.

Metagenomics changed that. By sequencing all the DNA in an environmental sample (a scoop of sediment, a liter of seawater, a gram of human stool), researchers can computationally reconstruct individual genomes from the mixed-up reads. These metagenome-assembled genomes, or MAGs, have opened up whole branches of the tree of life that were invisible before.15PubMed Central. Metagenome-Assembled Genomes (MAGs): Advances, Challenges, and Ecological Insights To bring some order to this flood of data, the Genomic Standards Consortium developed minimum reporting standards for MAGs, covering assembly quality, estimated completeness, and contamination levels.16Nature Biotechnology. Minimum information about a single amplified genome (MISAG) and a metagenome-assembled genome (MIMAG) of bacteria and archaea

Naming these organisms presents its own headache. The formal rules governing prokaryotic nomenclature still require a physical type specimen, which is impossible for something that has never been cultured. The provisional label “Candidatus” is used for such taxa, but it exists in a regulatory gray area: it is not formally recognized in the naming code, and proposals to allow gene sequences as type specimens have been rejected.17PubMed Central. Nomenclature of prokaryotic ‘Candidatus’ taxa: establishing order in the current chaos The result is a strange limbo where entire phyla of organisms have been described genomically but lack valid names under the official rules.

The OTU Versus ASV Debate

When researchers sequence 16S genes from an environmental sample, they end up with thousands or millions of short DNA reads that need to be grouped into meaningful units. For years, the standard approach was to cluster similar sequences into operational taxonomic units (OTUs), typically using a 97% similarity cutoff as a rough proxy for species. More recently, computational methods that resolve individual sequence variants (ASVs) without any clustering threshold have gained popularity. The idea is that ASVs offer finer resolution since they preserve every real sequence difference in the data.

The tradeoff is not as clean as it sounds. Because many bacteria carry multiple copies of the 16S gene, and those copies are not always identical within a single genome, ASVs can split one organism into several apparent taxa. An analysis of over 20,000 bacterial genomes found that as the number of 16S gene copies in a genome increased, so did the number of distinct ASVs. For a bacterium with seven copies, like E. coli, a distance threshold of over 5% was needed to keep all copies in a single cluster with 95% confidence.18PubMed Central. Amplicon Sequence Variants Artificially Split Bacterial Genomes into Separate Clusters In practical terms, ASV-based analyses produce fewer total units per sample than OTU-based ones, but the choice between methods has a stronger effect on diversity measurements than other analytical decisions like rarefaction depth.19PLOS ONE. Ranking the biases: The choice of OTUs vs. ASVs in 16S rRNA amplicon data analysis has stronger effects on diversity measures than rarefaction and OTU identity threshold

Why Horizontal Gene Transfer Blurs the Lines

Drawing a neat family tree for microbes runs into a fundamental biological problem: bacteria and archaea swap genes sideways, not just vertically from parent to offspring. Horizontal gene transfer is pervasive in the microbial world, and it means that different genes in the same genome can tell different evolutionary stories.20PubMed Central. Horizontal Gene Transfer and the History of Life One gene might place a bacterium close to a soil-dwelling relative; another gene might link it to an ocean-dwelling species that donated DNA through a virus or a shared environment.

This does not make classification hopeless, but it does mean researchers have to think carefully about what a “tree” represents when branches can exchange genetic material. Modeling work suggests that standard tree-building methods are reasonably robust to gene transfer, maintaining good performance even when individual gene trees disagree with each other substantially. Gene tree disagreement alone is not proof that transfer happened; other processes like incomplete sorting of ancestral variation can produce similar patterns.21Systematic Biology. A Model of Horizontal Gene Transfer and the Bacterial Phylogeny Problem Still, gene transfer has thrown the concept of a single, clean tree of life into a more complicated light, and methods that can account for reticulate (network-like) evolution remain an active area of development.

Virus Classification Plays by Different Rules

Viruses are not cells. They lack ribosomes, so there is no 16S or ITS gene to sequence. They do not share a single gene that is universal across all viruses the way ribosomal genes are universal across cellular life. This forced virologists to build an entirely separate classification system.

The most enduring framework is the Baltimore classification, published in 1971, which groups viruses by how they store and express their genetic information. The original six classes (expanded to seven) cover every known strategy: double-stranded DNA, single-stranded DNA, double-stranded RNA, positive-sense single-stranded RNA, negative-sense single-stranded RNA, RNA viruses that reverse-transcribe, and DNA viruses that reverse-transcribe.22PubMed Central. The Baltimore Classification of Viruses 50 Years Later: How Does It Stand in the Light of Virus Evolution? These classes remain the conceptual backbone of virology, even as genomic data has massively expanded the known diversity of viruses.

On top of the Baltimore system, the International Committee on Taxonomy of Viruses (ICTV) has built a formal megataxonomy organized into six realms, analogous to domains in cellular life. These realms are assembled based on shared hallmark proteins involved in building viral shells or copying viral genomes, and they are intended to reflect evolutionary relationships.23The ISME Journal. Megataxonomy and global ecology of the virosphere The explosion of viruses discovered through metagenomics has pushed this taxonomy to accommodate sequences found in environmental samples, though quality control of sequence data remains an ongoing concern.24PLOS Biology. Four principles to establish a universal virus taxonomy

The Protist Problem

Protists, the grab-bag of eukaryotic microbes that are not animals, plants, or fungi, present some of the thorniest classification challenges. Early molecular phylogenies sorted eukaryotes into a handful of “supergroups” with names like Chromalveolata, Excavata, and Rhizaria. Those groupings have been repeatedly reshuffled as more data arrived. Phylogenomic analyses showed that the supergroup Chromalveolata, which was supposed to unite organisms sharing a plastid derived from red algae, does not hold together as a natural group.25PLOS ONE. Phylogenomics Reshuffles the Eukaryotic Supergroups

A revised classification published in 2012 attempted to stabilize the nomenclature while acknowledging that deep nodes in the eukaryotic tree remained statistically unresolved.26PubMed Central. The revised classification of eukaryotes Since then, further work has continued to shuffle lineages, and most of the original supergroups have either been absorbed into new taxa or dissolved altogether.27Trends in Ecology & Evolution. How Are Microbes Classified? From Traits to Genetics The base of the eukaryotic tree remains one of the hardest parts of the tree of life to pin down, partly because the events happened so long ago and partly because endosymbiosis and gene transfer have tangled the signals. Recent analysis of Asgard archaea, the closest known relatives of eukaryotes, found that most conserved eukaryotic functional systems trace back to these archaea, with a more limited contribution from the bacterial ancestor of mitochondria.28PubMed Central. Dominant contribution of Asgard archaea to eukaryogenesis

When Species Boundaries Get Fuzzy

Even with genome-scale data, deciding where one microbial species ends and another begins can be genuinely ambiguous. The cyanobacterium Microcystis, a common bloom-forming organism in lakes, illustrates this well. Morphology-based classification recognizes multiple species, distinguished mainly by colony shape and cell arrangement. But DNA-based methods collapse most of those morphospecies into a single species with multiple ecotypes. A pangenome analysis of 122 Microcystis genomes found at least 16 genetically distinct groups that qualify as separate species by genomic criteria, most of which contain organisms previously called Microcystis aeruginosa based on appearance.29PubMed Central. Microcystis pangenome reveals cryptic diversity within and across morphospecies Morphology underestimates diversity in some directions and overestimates it in others.

Pangenome approaches, which consider not just the genes shared by all members of a group but also the genes found in only some members, are becoming important for capturing this fine-scale variation. New tools are being developed that use pangenome graphs to classify metagenomic reads down to the strain level, a resolution that single marker genes cannot reach.30bioRxiv. PanTax: Strain-level taxonomic classification of metagenomic data using pangenome graphs This is the frontier of microbial classification: moving beyond “what species is this” to “which strain is this, and what can it do.”

Machine Learning Enters the Picture

The sheer volume of sequence data generated by modern metagenomics has made computational classification a bottleneck. Traditional approaches compare new sequences against databases using alignment algorithms, which becomes slower as databases grow. Machine learning and deep learning models have been adapted to handle taxonomic assignment, and in some cases they match or exceed the accuracy of alignment-based tools.31PubMed Central. Machine Learning and Deep Learning Applications in Metagenomic Taxonomy and Functional Annotation The catch is robustness: these models tend to be trained on particular environments or database snapshots, and their performance on novel environments or newly sequenced organisms is less predictable. Databases grow fast, and a model trained on last year’s reference set may miss lineages added since. For now, machine learning acts as a powerful accelerator for classification, but it supplements rather than replaces the underlying genomic framework that defines what a microbial taxon is.