Tajima’s D: A Test for Natural Selection

Tajima’s D is a statistical test that compares two different ways of measuring genetic variation within a population to detect whether natural selection, rather than random drift alone, has been shaping that variation. Developed by geneticist Fumio Tajima in 1989, the test works by asking a deceptively simple question: do the number of mutations in a DNA sample and the frequency at which those mutations appear tell the same story? When those two estimates agree, evolution looks neutral. When they disagree, something interesting is happening, and that something could be selection, demographic upheaval, or both.

What a Positive or Negative Value Means

Tajima’s D compares two estimates of genetic diversity in a sample of DNA sequences. One estimate counts how many variable sites exist. The other weights those sites by how common each variant actually is in the population. When both estimates are roughly equal, D hovers near zero, consistent with the idea that mutations are accumulating and drifting without any particular push from selection. The test’s power comes from what happens when those estimates pull apart.

A negative D means there is an excess of rare variants in the sample. Picture a stretch of DNA where most individuals carry the same sequence, but a handful of individuals each carry their own unique mutation. That pattern produces more variable sites than you would expect given the overall level of pairwise differences, dragging D below zero. A positive D means the opposite: variants are hanging around at intermediate frequencies instead of being rare. Many individuals carry one version, many carry the other, and there are fewer rare oddball mutations than you would predict from the total count of variable sites.

Each direction points to different evolutionary stories. A negative value is consistent with recent positive selection (a “selective sweep” that wiped out nearby variation and left behind a crop of new, rare mutations), purifying selection (harmful mutations kept at low frequency), or rapid population growth (which floods the gene pool with fresh, rare variants). A positive value is consistent with balancing selection (where having two different versions of a gene is advantageous, keeping both at intermediate frequency) or a population contraction that eliminated rare variants and left behind the common ones.

The Demography Problem

The biggest challenge with Tajima’s D is that population history and natural selection can produce identical signatures. A population that recently exploded in size generates a wave of new mutations, most of which are rare, just like a selective sweep does. A population that went through a bottleneck loses rare variants and retains common ones, looking a lot like balancing selection. This means a single D value at a single gene, taken in isolation, cannot tell you whether selection or demographics is responsible.

A study of 151 genetic loci in African-American and European-American populations illustrates this clearly. The average Tajima’s D across the genome was negative in the African-American sample (around −0.49) and positive in the European-American sample (around +0.26). Those values look like they could reflect selection acting in opposite directions. But coalescent simulations showed that the negative values in African-Americans were well explained by population admixture, while the positive values in European-Americans fit a bottleneck model consistent with the demographic history of migration out of Africa. The genome-wide pattern in European-Americans showed a significant deviation from neutral expectations, consistent with a recent population contraction rather than widespread balancing selection.1Molecular Biology and Evolution. Disentangling the Effects of Demography and Selection in Human History

Population structure introduces its own artifacts. In teosinte, the wild ancestor of maize, Tajima’s D values calculated from species-wide samples were systematically more negative than values from samples taken within individual subpopulations. The subpopulation-specific estimates were closer to zero, more closely matching neutral expectations, because they removed the artificial excess of rare variants created by lumping genetically distinct groups together.2PubMed Central. Population structure and its effects on patterns of nucleotide polymorphism in teosinte (Zea mays ssp. parviglumis)

Researchers handle this ambiguity in a few ways. The most common approach is to compare the D value at a gene of interest against the genome-wide average. Demography affects all loci across the genome roughly equally, while selection targets specific regions. So a gene with a D value far more extreme than the genomic background becomes a candidate for selection even when the overall genome shows a demographic skew. Another strategy involves fitting demographic models first and then testing individual loci against expectations under that fitted model, using simulations to ask whether the observed D is more extreme than what the estimated demographic history alone would produce.3Molecular Biology and Evolution. Inferring the Effects of Demography and Selection on Drosophila melanogaster Populations from a Chromosome-Wide Scan of DNA Variation

Purifying Selection and Slightly Harmful Mutations

One of the clearest real-world applications of Tajima’s D involves detecting purifying selection, the quiet, constant process by which harmful mutations are weeded out of a population. The logic is straightforward: if a mutation damages a protein, it stays rare because the individuals carrying it leave fewer descendants. That keeps a steady supply of rare variants at functionally important sites, pulling D negative.

A large survey of bacterial genomes provided strong evidence for this. At synonymous sites, which do not change the protein product of a gene, Tajima’s D was positive in about 70% of datasets, with an overall median of roughly +0.87. At nonsynonymous sites, which do alter the protein, D was negative in about 68% of datasets, with a median around −0.66. The contrast is striking: mutations that change protein function are systematically kept rare, while those that leave the protein untouched drift to higher frequencies. This pattern is exactly what the nearly neutral theory predicts for populations with large effective sizes, where even slightly harmful mutations get efficiently removed.4PubMed Central. Evidence for abundant slightly deleterious polymorphisms in bacterial populations

The same signal has been documented in eukaryotic parasites. In the intestinal parasite Giardia duodenalis, significantly negative Tajima’s D values at nonsynonymous sites in a surface protein gene confirmed ongoing removal of slightly deleterious mutations, again consistent with the nearly neutral theory’s predictions.5PubMed. Clonal diversity in Giardia duodenalis isolates from Thailand: evidences for intragenic recombination and purifying selection at the beta giardin locus

Simulations have shown that the strength of purifying selection matters for how the signal appears. Weakly harmful mutations barely reduce overall diversity but still produce a noticeable skew toward rare variants. Moderately harmful mutations cause the largest reduction in diversity and the most dramatic distortion in variant frequencies, making them the easiest to detect with Tajima’s D. Very strongly harmful mutations are eliminated so quickly they barely register as polymorphisms at all.6Genetics. Toward an Evolutionarily Appropriate Null Model: Jointly Inferring Demography and Purifying Selection

Balancing Selection and Elevated Diversity

On the other end of the spectrum, Tajima’s D can flag genes where natural selection is actively maintaining more than one variant, a process called balancing selection. When two versions of a gene each offer an advantage under different circumstances, such as resistance to different strains of a pathogen, selection keeps both versions circulating at intermediate frequencies rather than allowing one to dominate. This inflates D because the proportion of intermediate-frequency variants is higher than expected under neutrality.

A genome-wide study of a planktonic crustacean, Daphnia, searched for genomic regions showing the combined hallmarks of balancing selection: elevated nucleotide diversity, elevated Tajima’s D, and reduced differentiation between populations. The researchers identified candidate regions for long-term pathogen-resistance genes where these signatures significantly exceeded the genomic background, regardless of the size of the sliding window used. The polymorphisms at these loci were shared across species that diverged millions of years ago, suggesting that balancing selection had been maintaining the same variants over deep evolutionary time.7Nature Communications. Long-term balancing selection for pathogen resistance maintains trans-species polymorphisms in a planktonic crustacean

The classic human example involves the immune system’s MHC genes, which present fragments of pathogens to immune cells. These genes consistently show elevated Tajima’s D values because having diverse versions helps an individual recognize a wider range of infections. The principle extends broadly: wherever maintaining genetic diversity is itself advantageous, you can expect D to be pushed positive at those specific loci.

Scanning Genomes for Selective Sweeps

When a beneficial mutation rises rapidly in frequency, it drags along nearby variants on the same chromosome, erasing diversity across a genomic neighborhood. After the sweep, new mutations begin accumulating again, but they are all recent and therefore rare. This produces a sharply negative Tajima’s D in the swept region, flanked by normal values on either side. Genome-wide scans exploit this pattern to find genes that have recently experienced strong positive selection.

An early genome scan using dense genotype data across the human genome identified contiguous regions of markedly reduced Tajima’s D. The study found that Tajima’s D from genome-wide genotyping data correlated significantly with D from full resequencing data across 179 genes, validating the approach. The scan turned up seven candidate sweep regions in an African-descent population, 23 in a European-descent population, and 29 in a Chinese-descent population.8PubMed Central. Genomic regions exhibiting positive selection identified from dense genotype data The larger number of sweeps detected outside Africa aligns with the expectation that populations adapting to new environments during and after the migration out of Africa would carry more signatures of recent positive selection.

The relationship between Tajima’s D and recombination rate adds another dimension to these scans. In the European-American sample from the 151-locus study, there was a significant positive relationship between recombination rate and Tajima’s D, even after controlling for the overall level of genetic variation. Regions with low recombination had more negative D values. This makes sense if selective sweeps have been occurring repeatedly across the genome: in low-recombination regions, the diversity-erasing effect of a sweep extends further along the chromosome, because nearby variants are less likely to be separated from the swept mutation by recombination. The pattern was consistent with multiple hitchhiking events associated with adaptation to novel habitats after the migration out of Africa.1Molecular Biology and Evolution. Disentangling the Effects of Demography and Selection in Human History

In domesticated crops, Tajima’s D works alongside other statistics to distinguish between “hard” and “soft” selective sweeps. A hard sweep starts from a single new mutation; a soft sweep involves selection favoring a variant that already existed at some frequency, or multiple independent mutations achieving the same effect. A study of over 300 wild and domesticated soybean accessions found that hard sweeps predominated in landraces and improved varieties, while soft sweeps were more prevalent in wild populations. The candidate loci identified overlapped with previously mapped regions linked to domestication traits like seed size and flowering time.9PubMed Central. Hard versus soft selective sweeps during domestication and improvement in soybean

Viral Evolution and Outbreak Monitoring

The speed and volume of viral evolution make Tajima’s D surprisingly useful for tracking pathogen dynamics. When a new viral variant emerges and rapidly spreads, the resulting population of viral genomes looks a lot like a selective sweep or a population expansion: lots of sequences that are very similar to one another, with variation mostly consisting of rare, recently acquired mutations. That drives D strongly negative.

During the COVID-19 pandemic, researchers applied Tajima’s D to whole-genome sequences of the Omicron variant and found values of about −2.7 across the full genome and −2.6 in the ORF1ab gene, both significantly below zero. The pattern indicated an excess of low-frequency variants consistent with strong selection or rapid demographic expansion. The authors argued that these low diversity and negative D values suggested Omicron had been spreading within a population for weeks before samples were collected, and that Tajima’s D could serve as an early warning system for emerging outbreaks.10medRxiv. Tajima D test accurately forecasts Omicron / COVID-19 outbreak

The test has also revealed shifts in evolutionary dynamics over time within a single virus lineage. Analysis of the spike gene in the common cold coronavirus HCoV-OC43 found positive Tajima’s D values during the pre-pandemic period, consistent with stable endemic circulation where multiple lineages coexist at intermediate frequencies. After the COVID-19 pandemic, D turned negative, suggesting the virus had undergone a population expansion following a bottleneck, likely caused by pandemic-era social distancing measures that disrupted its normal transmission patterns.11Enfermedades Infecciosas y Microbiología Clínica. Spike gene evolution in HCoV-OC43: Evidence of post-pandemic selective expansion

Practical Pitfalls in Modern Genomic Data

Tajima’s D was developed in an era when researchers sequenced a handful of gene copies by hand. Modern sequencing generates millions of short reads with variable coverage across the genome, and this creates systematic biases that many users do not realize are there.

Low sequencing depth is a major culprit. When coverage is shallow, rare variants are more likely to be missed entirely because they simply are not represented in enough sequencing reads to be confidently identified. This biases the frequency spectrum toward common variants, inflating Tajima’s D. The bias is not uniform across a genome; coverage varies from region to region depending on DNA content, repeat structure, and library preparation. That means regional comparisons of D values can be confounded by coverage differences rather than reflecting actual evolutionary differences.12PubMed Central. Calculation of Tajima’s D and other neutrality test statistics from low depth next-generation sequencing data

Missing genotype data introduces additional problems, and the direction of the bias depends on which software you use. A systematic evaluation found that as missing genotypes increased, three widely used tools (VCFtools, PopGenome, and scikit-allel) all overestimated Tajima’s D, with the strongest effect in VCFtools. A fourth tool, pegas, showed the opposite bias, underestimating D as missing data increased. Missing sites, by contrast, had no significant effect in any of the programs tested.13PubMed Central. Correcting for Bias in Estimates of θw and Tajima’s D From Missing Data in Next‐Generation Sequencing For researchers running genome scans where hundreds of thousands of windows are compared, even a small systematic bias tied to data quality can produce false signals of selection or mask real ones.

Window size is another decision that shapes results. In genome scans, Tajima’s D is typically calculated across sliding windows of a fixed number of base pairs. Too small a window captures few variable sites and produces noisy estimates; too large a window averages over regions with very different evolutionary histories. The bat migration study mentioned earlier, for example, used 4,000-base-pair windows when scanning for candidate regions associated with migratory behavior.14Genome Biology and Evolution. To Migrate or not to Migrate? Exploring the Genomic Basis of Partial Migratory Behavior in Bats Choosing the right window size is part art, part informed guesswork, and part sensitivity analysis, and different window sizes can highlight different candidate regions for the same dataset.

When Tajima’s D Is Used Alongside Other Tests

No experienced researcher relies on Tajima’s D alone to conclude that selection is acting on a gene. The test belongs to a family of related statistics, each of which slices the frequency spectrum differently. Fu and Li’s D and F statistics weight the comparison differently, emphasizing mutations on external versus internal branches of the genealogy. Fay and Wu’s H is specifically sensitive to high-frequency derived variants, which are a hallmark of hitchhiking during a selective sweep. A class of tests including Tajima’s D as well as the Fu and Li statistics has been formally evaluated for statistical properties like power and false-positive rates, and all perform differently depending on the type of selection and the demographic scenario involved.15Oxford Academic (Genetics). Properties of statistical tests of neutrality for DNA polymorphism data

In practice, researchers layer multiple approaches. A region showing strongly negative Tajima’s D, extended haplotype homozygosity (long stretches of identical sequence, indicating a recent sweep), and elevated population differentiation between groups becomes a far more compelling candidate for positive selection than any one signal alone. Population-genetic estimators have also been adapted for pooled sequencing approaches, where DNA from many individuals is sequenced together to save costs, with corrections built in for the different error profiles of pool-seq data.16PubMed Central. Population genomics from pool sequencing The soybean domestication study, for instance, identified its strongest candidate loci by requiring overlap between differentiation-based statistics and haplotype-based statistics, filtering out the many false positives that any single test would have generated.

The underlying reality is that Tajima’s D is a first-pass screening tool. It flags regions of the genome where the pattern of variation departs from simple neutral expectations. Figuring out why it departs requires additional evidence: functional data about what the gene does, comparisons across populations with different demographic histories, and ideally experimental or epidemiological evidence that the variant in question actually affects fitness. The test opened a door in 1989 that remains heavily trafficked, but the rooms behind it require more than one key to explore.