How Can DNA Be Used to Classify Organisms?

DNA classifies organisms by serving as a molecular identity card: researchers read short, standardized stretches of an organism’s genetic code and compare those sequences against databases of known species. When two samples share a very high percentage of matching code in the right gene region, they belong to the same species or a closely related one. When the sequences diverge beyond a certain threshold, the organisms are classified as distinct. This approach has reshaped biology over the past two decades, revealing hidden species that look identical to the eye and correcting long-standing misidentifications based on physical appearance alone.

DNA Barcoding and the Idea of a Genetic Name Tag

The most widespread method for identifying organisms by DNA is called barcoding. The concept is straightforward: just as a barcode on a grocery item points to one specific product, a short stretch of DNA can point to one specific species. For most animals, that stretch comes from a mitochondrial gene called cytochrome c oxidase subunit I, commonly abbreviated COI. Because mitochondria are passed from mother to offspring with relatively little shuffling, COI sequences tend to be consistent within a species and different enough between species to tell them apart.1PubMed Central. Is There a Key Primer for Amplification of Core Land Plant DNA Barcode Regions (rbcL and matK)?

Plants, however, have slower mutation rates in their mitochondrial DNA, so the COI gene does not vary enough to be useful. Instead, plant barcoding relies on two genes in the chloroplast genome, rbcL and matK, which were recommended as core barcode regions by an international working group.1PubMed Central. Is There a Key Primer for Amplification of Core Land Plant DNA Barcode Regions (rbcL and matK)? A third marker, ITS, from the nuclear genome, is sometimes added to improve resolution. A study of the cashew family (Anacardiaceae) found that combining matK, rbcL, and ITS together was the most effective barcode for correctly grouping species within the same genus and separating unrelated ones.2Jurnal Biologi Tropis. Analysis of Three DNA Barcoding (matK, rbcL, ITS) in the Anacardiaceae Family

The process itself is remarkably simple in outline. A researcher collects a tissue sample, extracts DNA, amplifies the target barcode region, sequences it, and then searches a reference database for the closest match. Two major databases serve this purpose: the Barcode of Life Data Systems (BOLD) and the National Center for Biotechnology Information (NCBI). Both have quality issues. An evaluation of marine species in the western and central Pacific found that NCBI had broader coverage but lower sequence quality, while BOLD was more curated but less complete. Problems in both databases included short or ambiguous sequences, incomplete taxonomic labels, and internal conflicts likely caused by contamination, sequencing errors, or cryptic species lumped under one name.3PubMed Central. Evaluation of DNA barcoding reference databases for marine species in the western and central Pacific Ocean Database quality matters because a barcode search is only as good as the library it searches against.

How Microbes Get Classified Without a Body to Examine

Bacteria and archaea do not have limbs, shells, or leaves to measure. For decades, microbiologists have relied on a single gene, the 16S ribosomal RNA gene, to classify these organisms. The gene is roughly 1,500 base pairs long and contains regions that are nearly identical across all bacteria (useful for designing universal primers to grab it) alongside regions that vary enough to distinguish one group from another.4PubMed Central. Taxonomic classification of bacterial 16S rRNA genes using short sequencing reads: evaluation of effective study designs Sequencing the 16S gene and comparing it against databases of known type strains allows researchers to place an unknown bacterium within the tree of life, from the broadest classification down to genus or species level.5PubMed Central. Setting new boundaries of 16S rRNA gene identity for prokaryotic taxonomy

For a long time, most studies sequenced only short fragments of the 16S gene, covering one or two of its variable regions. That was sufficient for sorting bacteria into broad groups but often too coarse to distinguish between closely related species. Sequencing the full-length gene provides much finer resolution. Research has shown that full-length 16S sequencing can achieve species-level and even strain-level identification when intragenomic variation between the multiple copies of the 16S gene within a single cell is properly accounted for.6Nature Communications. Evaluation of 16S rRNA gene sequencing for species and strain-level microbiome analysis Newer long-read sequencing platforms from PacBio and Oxford Nanopore have made full-length 16S sequencing practical and affordable, pushing microbial identification closer to the species level in routine environmental and clinical studies.7PubMed Central. Full-length 16S rRNA gene sequencing by PacBio improves taxonomic resolution in human microbiome samples8PubMed Central. The newest Oxford Nanopore R10.4.1 full-length 16S rRNA sequencing enables the accurate resolution of species-level microbial community profiling

Even with full-length sequencing, though, a persistent question nags at the field: what percentage of sequence similarity should count as the boundary between species? For years, a 97% similarity threshold was widely used to cluster sequences into operational units meant to approximate species. That threshold was always a rough convenience. Studies have shown that its reliability breaks down when applied to short amplicons covering only one or two variable regions, and that in some groups the cutoff lumps genuinely distinct species together while in others it splits a single species apart.9PubMed Central. Reconciliation between operational taxonomic units and species boundaries10PubMed. Molecular operational taxonomic units as approximations of species in the light of evolutionary models and empirical data from Fungi There is growing consensus that microorganism taxonomy needs to shift more formally toward DNA-based approaches, including updated thresholds calibrated against larger collections of type strains.11PubMed. Toward DNA-based taxonomy of prokaryotes and microeukaryotes

Whole Genomes and Phylogenomics

Single-gene barcoding works well for quick identifications, but it has limits. A single gene captures only a tiny slice of an organism’s evolutionary history. When researchers want to reconstruct deeper relationships, resolve ambiguous cases, or classify bacteria below the species level, they turn to whole-genome approaches. Comparing entire genomes provides far more data points and far higher resolution than any single marker can offer.12PubMed Central. Whole-Genome Sequencing of Bacterial Pathogens: the Future of Nosocomial Outbreak Analysis

In microbiology, whole-genome sequencing (WGS) has become the gold standard for precise classification. One technique, digital DNA-DNA hybridization, computationally compares two entire genomes and expresses their similarity as a percentage. A threshold of 70% is the traditional cutoff for declaring two organisms the same species. In a study of waterborne Enterobacter isolates from South Africa, researchers found that their unknown samples matched the subspecies Enterobacter hormaechei subsp. hoffmannii with digital hybridization values above 92%, well above the species cutoff, confirming the identification with high confidence.13PubMed Central. Whole Genome Sequencing Based Taxonomic Classification, and Comparative Genomic Analysis of Potentially Human Pathogenic Enterobacter spp. Isolated from Chlorinated Wastewater in the North West Province, South Africa

For plants, animals, and fungi, the genome-scale approach is called phylogenomics: building evolutionary trees from hundreds or thousands of genes simultaneously rather than one or two. High-throughput sequencing has made this feasible even for non-model organisms, and pipelines now exist to extract useful phylogenetic markers from low-coverage genome data without needing a polished reference genome.14Genome Biology and Evolution. A Practical Guide to Design and Assess a Phylogenomic Study15Methods in Ecology and Evolution. Phylogenomics from low‐coverage whole‐genome sequencing Researchers align the sequences, build trees using methods like maximum likelihood or Bayesian inference, and look for the branching pattern that best explains the observed genetic differences.16PubMed Central. Common Methods for Phylogenetic Tree Construction and Their Implementation in R

When Physical Appearance Gets It Wrong

One of DNA’s most dramatic contributions to classification has been revealing cryptic species, organisms that are genetically distinct but physically indistinguishable. A study of freshwater snails in the genus Radix found that shell shape was essentially useless for telling species apart, while DNA sequences cleanly separated them. The authors argued that this situation is widespread among invertebrates and proposed DNA taxonomy as a more reliable and objective identification method.17PubMed Central. Comparing the efficacy of morphologic and DNA-based taxonomy in the freshwater gastropod genus Radix (Basommatophora, Pulmonata)

In fish, similar discoveries are common. A study of the Japanese cyprinid Biwia zezera found two deeply divergent genetic groups with about 8.6% sequence difference in the mitochondrial cytochrome b gene, confirmed by three independent nuclear DNA markers. The population in the Yodo River system turned out to be an entirely undescribed species hiding in plain sight.18PubMed. Population divergence of Biwia zezera (Cyprinidae: Gobioninae) and the discovery of a cryptic species, based on mitochondrial and nuclear DNA sequence analyses Lampreys tell a similar story: two cryptic species of Lethenteron that could not be distinguished by appearance showed about 9.1% sequence divergence in their COI gene, and researchers developed a simple PCR test using species-specific primers that could identify either species from a tissue sample.19Journal of Fish Biology. Mitochondrial DNA sequence divergence between two cryptic species of Lethenteron, with reference to an improved identification technique

Cryptic diversity matters beyond academic cataloging. Conservation programs that lump two distinct species under one name may unknowingly let one of them decline toward extinction. Fisheries management, invasive-species policy, and habitat protection all depend on knowing exactly what species are present. DNA-based identification has become the practical tool that turns “one species” into “actually, two.”

Reading DNA From Water and Soil

You do not always need an organism in hand to identify it. Environmental DNA, or eDNA, refers to the genetic material that organisms shed into their surroundings through skin cells, mucus, feces, pollen, or root exudates. By filtering a water sample and extracting the DNA from whatever is floating in it, researchers can identify which species are present in a river, lake, or ocean without ever seeing or catching them.20PubMed Central. Environmental DNA Metabarcoding: A Novel Contrivance for Documenting Terrestrial Biodiversity

The technique is called eDNA metabarcoding when it targets many species at once. Instead of designing primers for a single target species, researchers use broad primers that amplify barcode regions from a wide range of taxa, then sequence everything in one run and match the results against reference databases. A recent study along the Upper Rhine River between Basel and Strasbourg used multi-marker eDNA metabarcoding to simultaneously survey aquatic plants, invasive species, and vegetation associated with different land uses along the catchment, all from river water samples.21Ecological Indicators. Environmental DNA metabarcoding for catchment-scale detection of aquatic plants, invasive species, and land-use indicators in a large river That kind of integrated snapshot would take a conventional survey team weeks or months of field work to approximate.

Practical Uses Beyond the Lab Bench

DNA-based classification has moved well beyond academic taxonomy into everyday enforcement, public health, and consumer protection.

In forensic wildlife cases, DNA barcoding has become a courtroom tool. When South African authorities intercepted suspicious biological specimens, sequencing the COI gene identified the species involved in the crimes.22PubMed. DNA barcoding as a tool for species identification in three forensic wildlife cases in South Africa In Pakistan, a shipment labeled “fish meat” was intercepted at a port and tested with fish-specific barcoding primers. The sequences matched the Indian flap-shelled turtle, a species protected under the Convention on International Trade in Endangered Species (CITES), at 99% identity. What looked like a routine seafood shipment was actually an attempt to smuggle an endangered reptile across borders.23Journal of Bioresource Management. Use of DNA Barcoding to Control the Illegal Wildlife Trade: A CITES Case Report from Pakistan

Food fraud is a quieter but arguably more widespread problem. Mislabeled fish, adulterated meat, and counterfeit herbal medicines are common in global supply chains. DNA barcoding has been described as a gold-standard method for detecting this kind of fraud because it can objectively determine which species a product came from, regardless of how it has been processed, cooked, or disguised.24PubMed Central. Application of DNA barcoding for ensuring food safety and quality

In hospitals, whole-genome sequencing has transformed how disease outbreaks are tracked. When several patients on the same ward develop the same bacterial infection, doctors need to know whether a single strain is spreading through the facility or whether each patient picked up an unrelated strain from the community. WGS can discriminate between nearly identical bacterial lineages, and modeling studies have developed thresholds for defining when isolates are close enough genetically to be part of the same outbreak.25PubMed Central. Defining genomic epidemiology thresholds for common-source bacterial outbreaks: a modelling study WGS is increasingly used not just to react to outbreaks after the fact but to monitor pathogens prospectively and predict traits like antibiotic resistance from the genome before lab susceptibility tests come back.26PubMed. Genome sequencing for prevention of health-care-associated bacterial infections

Why DNA Classification Sometimes Gives Conflicting Answers

DNA-based classification is powerful, but it is not infallible. Several biological realities can cause different genes within the same organism to tell different evolutionary stories.

Horizontal gene transfer is the most disruptive. Bacteria routinely swap chunks of DNA with unrelated species, which means a gene acquired from a distant relative may suggest a close relationship that does not actually exist. Even rare transfers can produce wildly misleading phylogenies because a single horizontal event can make two gene trees completely incongruent.27PubMed Central. Horizontal Gene Transfer and the History of Life28Current Opinion in Microbiology. Horizontal gene transfer and phylogenetics This problem is severe enough in prokaryotes that the very concept of a single tree of life has been questioned.

Animals and plants face a different problem called incomplete lineage sorting. When species split apart recently, their ancestral gene variants may not have had time to sort cleanly into the new lineages. The result is that a gene tree built from one locus may place species A closer to species C, while a gene tree from a different locus places species A closer to species B.29PubMed. Coalescent-based species tree inference from gene tree topologies under incomplete lineage sorting by maximum likelihood This is one reason single-gene barcoding sometimes fails, and it is a major motivation for moving toward multi-gene and whole-genome approaches.

A related complication, called mitonuclear discordance, occurs when the mitochondrial genome and the nuclear genome point to different species boundaries. This can happen through ancient hybridization events where one species captures the mitochondria of another, or through incomplete lineage sorting at the mitochondrial level. In wolf spiders, for example, genomic analysis revealed that mitonuclear discordance significantly complicated species identification.30PubMed. Mitonuclear discordance in wolf spiders: Genomic evidence for species integrity and introgression The same phenomenon has been documented in rotifers, where mitochondrial barcodes and nuclear markers gave conflicting signals across multiple species complexes.31Zoologica Scripta. Mitonuclear discordance as a confounding factor in the DNA taxonomy of monogonont rotifers The practical takeaway is that relying on a single mitochondrial barcode alone can lead to both false lumping and false splitting. Multi-locus approaches and statistical methods designed to account for gene tree discordance are the standard safeguard.32PubMed Central. Phylogenomic species tree estimation in the presence of incomplete lineage sorting and horizontal gene transfer

Ancient DNA and the Classification of Extinct Life

DNA does not only classify living organisms. Over the past two decades, ancient DNA recovered from fossils, permafrost, and cave sediments has extended molecular classification to species that vanished thousands or even millions of years ago. Researchers have reconstructed genomes from extinct species and the ecosystems they inhabited, fundamentally changing our understanding of past biodiversity and evolutionary relationships.33PubMed. Ancient RNA expression profiles from the extinct woolly mammoth Ancient DNA work has revealed, for instance, that some organisms classified as single species based on fossil bones were genetically distinct populations, while others thought to be separate species were not. The frontier is now moving beyond DNA: researchers have recently recovered ancient RNA expression profiles from a preserved woolly mammoth, opening the possibility of understanding not just what genes extinct organisms carried but which ones were actually active.

Who Owns the Sequences

As DNA-based classification generates ever-larger databases of genetic sequences, a legal question has grown in parallel: who benefits when those sequences are used? The Convention on Biological Diversity and the Nagoya Protocol require benefit-sharing with the countries or communities that provide genetic resources. But “digital sequence information,” the nucleotide data uploaded to public databases, exists in a legal gray zone. Whether sharing a DNA sequence online constitutes sharing a genetic resource, triggering benefit-sharing obligations, remains unsettled under international law.34PubMed Central. Digital Sequence Information and the Access and Benefit-Sharing Obligation of the Convention on Biological Diversity

This is not abstract policy. Plant genebanks, which store seeds and tissue from thousands of species, increasingly generate and share DNA sequence data as part of their operations. How definitions of digital sequence information are drawn will determine whether uploading a barcode sequence to a public database triggers the same access and benefit-sharing obligations as physically sending a seed sample to another country.35PLANTS, PEOPLE, PLANET. Practical consequences of digital sequence information (DSI) definitions and access and benefit‐sharing scenarios from a plant genebank’s perspective If the rules become too restrictive, researchers worry that data sharing will slow to a trickle, undermining the reference databases that make DNA classification work in the first place. If the rules remain too loose, biodiversity-rich countries may see their genetic heritage cataloged and commercialized without compensation. The debate is ongoing and consequential for every branch of DNA-based taxonomy.