What Is Maximum Parsimony in Phylogenetic Analysis?

Maximum parsimony is a method for building evolutionary trees that works on a disarmingly simple principle: the best tree is the one that requires the fewest evolutionary changes to explain the observed data. If you have DNA sequences or physical traits from a set of species, maximum parsimony evaluates every possible tree arrangement and favors the one where the total number of mutations or character-state changes is smallest. The idea borrows from the broader philosophical concept of parsimony, sometimes called Occam’s razor, but applies it with mathematical rigor to the problem of figuring out how organisms are related. Despite being one of the oldest approaches in phylogenetics, it remains surprisingly relevant and is the subject of active research and debate.

The Core Logic

Imagine you have genetic sequences from five species and you want to know which ones share the most recent common ancestor. Each possible branching pattern, or tree topology, tells a different story about their history. Maximum parsimony scores each of those stories by counting the minimum number of changes (say, a nucleotide switching from A to G) needed along the branches to produce the sequences you actually observe at the tips. The tree that needs the fewest total changes wins. This idea applies equally to molecular data like DNA and to physical characteristics like tooth shape or limb structure.

At each internal node of the tree, the method infers what the ancestral state probably was. If two descendant lineages both carry the same nucleotide at a given position, the simplest explanation is that their ancestor carried it too. When descendants disagree, a change must have happened somewhere along the way. The method assigns ancestral states that minimize the total count of such changes across the entire tree.1PubMed. Quantifying the accuracy of ancestral state prediction in a phylogenetic tree under maximum parsimony This makes maximum parsimony a “character-based” approach: it works directly from the raw data at each site or trait, rather than first converting the data into a distance matrix.2PubMed Central. Maximum Parsimony on Phylogenetic networks

How the Counting Actually Gets Done

For a single tree with a known shape, calculating the parsimony score is computationally straightforward. Two classic algorithms handle it. The Fitch algorithm works for cases where any change between character states costs the same (an A-to-G mutation is no worse than an A-to-T mutation). It sweeps from the tips of the tree inward, assigning sets of possible states at each internal node, then sweeps back outward to pick the most parsimonious assignments. The Sankoff algorithm generalizes this by allowing different costs for different types of changes, which is useful when, for example, transitions between chemically similar nucleotides are known to happen more often than transversions between dissimilar ones. Both algorithms have been extended to work on phylogenetic networks, not just strictly branching trees, though reticulate vertices (where lineages merge, as in hybridization) introduce complications.2PubMed Central. Maximum Parsimony on Phylogenetic networks Dynamic programming approaches can solve these network-extended problems exactly, though at greater computational cost.3PubMed Central. Exactly computing the parsimony scores on phylogenetic networks using dynamic programming

The real computational challenge is not scoring a single tree but finding the best tree among an astronomical number of possibilities. For just 10 species, there are over two million possible unrooted tree topologies. For 50 species, the number exceeds anything that could be enumerated in the lifetime of the universe. This means that for anything beyond a handful of taxa, finding the globally most parsimonious tree is an exercise in heuristic searching, not exhaustive comparison.

Searching for the Best Tree

Because you cannot check every possible tree, parsimony analyses rely on clever search strategies that explore “tree space” efficiently. A typical search starts by building a rough initial tree, often using a method called Wagner addition, which adds taxa one at a time to whichever position produces the fewest extra steps. That starting tree is then improved through branch-swapping operations: the software breaks the tree apart and reconnects it in different configurations, keeping any rearrangement that lowers the parsimony score. The most thorough of these operations is called tree bisection-reconnection (TBR), which cuts the tree in two, tries every possible way to rejoin the halves, and keeps the best result.

The trouble is that these hill-climbing strategies can get stuck. Tree space is not a smooth landscape with one valley leading downhill to the optimal tree. It is rugged, full of local optima where every small rearrangement makes the score worse, even though a very different tree topology somewhere else in the landscape might be better overall. Research has shown that the shape of this landscape, not just the amount of conflicting signal in the data, is what makes some datasets so hard to solve. In some cases, even datasets with relatively little character conflict can present extremely rugged terrain, with narrow peaks and valleys that trap standard search algorithms.4PubMed. Hide and vanish: data sets where the most parsimonious tree is known but hard to find, and their implications for tree search methods Missing data makes things worse by creating flat plateaus where many very different trees share the same score, and the search algorithm wanders without making progress.

One influential solution to this problem is the parsimony ratchet, developed by Kevin Nixon. The ratchet works by periodically reweighting a random subset of characters, performing branch swapping on this perturbed dataset, then restoring the original weights and swapping again. This cycle of perturbation and recovery jolts the search out of local optima and into new regions of tree space. Tests on large datasets showed the ratchet finding good trees roughly 20 to 80 times faster than traditional approaches, and thousands of times faster than naive searches.5PubMed. The Parsimony Ratchet, a New Method for Rapid Parsimony Analysis Other strategies, such as tree drifting and sectorial searches, have also been developed and are available in modern software. These techniques allow parsimony analyses to handle datasets with hundreds or even thousands of species.6PubMed Central. Efficient tree searches with available algorithms

The Homoplasy Problem

Maximum parsimony works beautifully when evolutionary changes are rare and each change is informative about relationships. In that situation, shared changes between species really do reflect shared ancestry, and the simplest explanation is the correct one. The method runs into trouble when the same change evolves independently in unrelated lineages. A bird and a bat both have wings, but not because they inherited wings from a common ancestor; flight evolved separately. In molecular data, this kind of convergence is called homoplasy, and it is unavoidable because there are only four nucleotide states at any given site in a DNA sequence. Given enough time, the same mutation will happen independently in different lineages just by chance.

Homoplasy introduces two kinds of errors. Small datasets suffer from random noise: there simply are not enough characters to distinguish real shared ancestry from coincidental similarity. Large datasets mostly overcome that problem, but they face a subtler one. When certain branches in the tree are very long (meaning a lot of evolution happened along them), those lineages accumulate so many changes that they start to resemble each other by chance. Maximum parsimony can be misled into grouping these long-branched lineages together, a well-known artifact called long-branch attraction.7Current Biology. Systematic errors in phylogenetic trees This is not just a parsimony problem. Long-branch attraction can also affect likelihood and Bayesian methods when sequence lengths are limited.8PubMed. Long-Branch Attraction in Species Tree Estimation: Inconsistency of Partitioned Likelihood and Topology-Based Summary Methods

How Maximum Parsimony Compares to Probabilistic Methods

The field of phylogenetics has shifted substantially toward maximum likelihood and Bayesian inference over the past few decades. Both of these approaches use explicit models of how sequences evolve: they estimate the probability of the observed data given a particular tree and a set of parameters describing mutation rates, base frequencies, and rate variation among sites. Parsimony does not assume such a model, which is both its greatest strength and its most frequently cited weakness.

The case against parsimony is usually built around simulation studies showing that when sequences evolve according to a known, relatively uniform model, likelihood methods recover the true tree more reliably, especially when branch lengths vary a lot. This is the long-branch attraction scenario, and it was a major driver of the shift toward model-based methods. But the case is not as clean as textbooks sometimes suggest. A well-known study published in Nature demonstrated that when evolutionary rates at individual sites shift differently across the tree, a condition common in real data, maximum likelihood and Bayesian methods can themselves become strongly biased and statistically inconsistent. Under those heterogeneous conditions, parsimony outperformed both probabilistic approaches across a wide range of scenarios tested.9PubMed. Performance of maximum parsimony and likelihood phylogenetics when evolution is heterogeneous

For morphological data, the comparison is even more interesting. A large survey of empirical morphological matrices from MorphoBank found that Bayesian inference under the standard model for morphology produced more unresolved trees and wider credibility intervals than parsimony. Trees from the two methods frequently disagreed, and the number of species in the matrix was the strongest predictor of disagreement. The study concluded that Bayesian inference was less precise than parsimony for the morphological datasets examined.10PubMed. Comparative evaluation of maximum parsimony and Bayesian phylogenetic reconstruction using empirical morphological data On the other hand, when tested against a well-supported reference phylogeny for insects using morphological characters, likelihood and Bayesian methods produced trees with higher resolution and precision than parsimony.11PubMed Central. Performance of tree-building methods using a morphological dataset and a well-supported Hexapoda phylogeny In many real-world cases, though, the different methods converge on broadly similar results. A recent analysis of fossil and living hamsters found that parsimony and Bayesian topologies were largely congruent, confirming the same major groupings.12PubMed Central. Phylogenetic relationships of Neogene hamsters (Mammalia, Rodentia, Cricetinae) revealed under Bayesian inference and maximum parsimony

The honest summary is that no single method wins everywhere. Parsimony is most vulnerable when branch lengths are highly unequal and evolution is roughly model-like. Likelihood methods are most vulnerable when the model is badly wrong about how rates vary. Researchers who care about getting the right answer tend to run multiple methods and pay close attention to where they agree and disagree.

Implied Weighting and Other Refinements

One of the more important developments in modern parsimony is implied weighting. The standard version of maximum parsimony treats every character equally. A tooth-shape character with lots of convergent changes gets the same vote as a highly conserved gene region that changes only once. Implied weighting addresses this by automatically downweighting characters that show more homoplasy on the current best tree. Characters that conflict with the tree a lot are treated as less reliable, while characters that fit cleanly carry more influence. The weights update dynamically as the tree changes during the search.13Palaeontology. Implied weighting and its utility in palaeontological datasets: a study using modelled phylogenetic matrices

This approach has become widely used in paleontology, where morphological matrices often contain high levels of homoplasy and missing data. A study across 70 morphological datasets found that downweighting homoplastic characters produced clear improvements in support values, measured by jackknife resampling.14Cladistics. Weighting against homoplasy improves phylogenetic analysis of morphological data sets Extended versions of implied weighting now allow researchers to weight different data partitions differently. For example, morphological characters can be weighted against homoplasy while molecular characters are held at equal weight, or entire genes can be weighted according to their average level of conflict.15PubMed. Extended implied weighting This flexibility makes parsimony more adaptable to the messy, mixed datasets that are increasingly common in systematics.

Measuring Confidence in the Tree

Any phylogenetic method needs a way to tell you how confident you should be in each branch of the result. In parsimony, one commonly used measure is the decay index (sometimes called Bremer support). For a given branch in the most parsimonious tree, the decay index asks: how many extra steps would it take before a tree that lacks this branch becomes as short as the shortest trees that include it? A branch with a decay index of 5 means you would need to accept a tree five steps longer before you could find one that dissolves that particular grouping. Higher values indicate stronger support.16PubMed. A chain is no stronger than its weakest link: double decay analysis of phylogenetic hypotheses

Bootstrap resampling is another standard approach. The method repeatedly resamples characters from the dataset (drawing new datasets of the same size, with replacement), runs a parsimony analysis on each resampled dataset, and reports how often each branch appears in the resulting trees. Branches that show up in, say, 95 out of 100 bootstrap replicates are considered well supported. Jackknife resampling works similarly but drops a fraction of the characters each time instead of resampling with replacement. These resampling methods are not unique to parsimony; they are used across phylogenetic approaches, but they originated in the parsimony framework and remain central to how parsimony results are evaluated.

Software for Parsimony Analysis

Several major software packages implement maximum parsimony, and they differ substantially in what they can handle. TNT (Tree analysis using New Technology) is widely regarded as the most powerful option for large-scale parsimony searches. It implements the ratchet, tree drifting, sectorial searches, and tree fusing, and can handle datasets with thousands of taxa. PAUP* (Phylogenetic Analysis Using Parsimony) has long been a workhorse of the field, with flexible search options and well-tested algorithms, though it lacks some of the advanced search strategies that TNT offers. For very large and complex datasets, this can make finding the most parsimonious tree difficult. MEGA, a popular general-purpose phylogenetics package, also includes parsimony, but evaluations have found significant drawbacks in its implementation: it sometimes fails to find the most parsimonious trees because it does not perform all possible branch-swapping rearrangements.17PubMed Central. Parsimony analysis of phylogenomic datasets (II): evaluation of PAUP*, MEGA and MPBoot

For researchers doing primarily molecular phylogenetics, parsimony now often plays a supporting role. Fast maximum-likelihood programs like RAxML and IQ-TREE use parsimony internally as a shortcut during tree searching: they generate starting trees via parsimony and use parsimony scores to filter candidate tree rearrangements before evaluating them under the full likelihood criterion.18PubMed Central. Evaluating Fast Maximum Likelihood-Based Phylogenetic Programs Using Empirical Phylogenomic Data Sets Even if the final tree is chosen by likelihood, parsimony is doing real work under the hood to make the search tractable.

Where Parsimony Still Thrives

Parsimony remains the dominant method in several research communities. In paleontology and morphological systematics, where data consist of physical traits scored from fossils and living organisms, model-based methods face a fundamental challenge: there is no consensus on what the correct model of morphological evolution should look like. The most commonly used model for morphological Bayesian analysis makes assumptions (like equal rates of change among all character states) that many researchers find unrealistic. Parsimony avoids these assumptions entirely, which is why many paleontologists prefer it. Combined with implied weighting and the powerful search algorithms in TNT, parsimony analyses of large morphological matrices remain competitive and widely published.

In cladistic philosophy, parsimony also holds a special status. Some systematists argue that parsimony is epistemologically preferable because it does not require assuming a particular model of evolution. It simply asks which tree explains the data with the fewest ad hoc hypotheses of change. This argument has roots in Karl Popper’s philosophy of science, where the best hypothesis is the one that is most falsifiable and requires the fewest auxiliary assumptions. The relationship between parsimony and Popperian falsificationism has been debated extensively in the systematics literature, with no final resolution, but it remains a live philosophical issue for researchers who care about what phylogenetic methods actually claim about the world.

Parsimony Beyond Traditional Phylogenetics

The parsimony framework has found applications well outside the traditional domain of species-level evolutionary trees. One striking recent example is in cancer biology. Researchers studying how tumors spread through the body have used lineage-tracing tools based on CRISPR-Cas9 to mark individual cancer cells with unique, heritable “barcodes.” As cells divide and accumulate edits, parsimony-based tree reconstruction can infer the branching history of tens of thousands of cells, revealing which clones metastasized, what routes they took through the body, and what gene-expression differences drove their invasive behavior.19PubMed Central. Single-cell lineages reveal the rates, routes, and drivers of metastasis in cancer xenografts The appeal of parsimony here is practical: when you have extremely large numbers of cells and relatively few informative characters (the barcode edits), parsimony is fast, interpretable, and requires no assumptions about mutation rates across lineages.

Similar logic applies to tracking the spread of viruses during outbreaks. Viral genomes are short, mutation rates can vary across lineages, and datasets can be enormous. Parsimony provides quick, rough phylogenies that epidemiologists use as a first pass to identify clusters and transmission chains, even if more refined analyses using likelihood or Bayesian methods follow later. In language evolution, too, parsimony has been used to study how words and grammatical features change across language families, treating linguistic traits much like morphological characters in a cladistic analysis. The versatility of the method comes from its simplicity: anywhere you have discrete characters that change over a branching process, parsimony can provide an initial reconstruction without requiring you to specify how fast or in what pattern those changes occur.