AncestryDNA vs 23andMe: Which Is More Accurate?

Neither AncestryDNA nor 23andMe is meaningfully more accurate than the other at reading your actual DNA. Both use genotyping chips that capture hundreds of thousands of genetic markers with high reliability, and both produce raw data of comparable technical quality. Where the two companies diverge, sometimes dramatically, is in how they interpret that data, particularly when estimating your ethnic origins. A study of identical twins found that when the same person’s DNA was tested by two different companies, ancestry-composition agreement dropped as low as roughly 53%, compared with near-perfect agreement when the same company tested both twins. That gap has less to do with laboratory precision and more to do with the statistical models, reference populations, and algorithms each company uses behind the scenes.

What “Accuracy” Means in Consumer DNA Testing

When people ask which test is more accurate, they usually mean one of three things: How reliably does the lab read my DNA? How close are the ethnicity percentages to my actual heritage? And how trustworthy are the health reports? These are genuinely different questions with different answers, and conflating them is where most confusion starts.

At the genotyping level, both companies use Illumina microarray chips to read specific positions across your genome. Saliva-based collection, the method both services use, produces DNA of sufficient quality for high-throughput genotyping, with studies showing genotype call rates and concordance above 97% when saliva samples are compared against blood-derived DNA.1PubMed Central. Saliva samples are a viable alternative to blood samples as a source of DNA for high throughput genotyping Separate research in a predominantly elderly population confirmed that Illumina platforms handle saliva DNA reliably even when the samples contain impurities or bacterial DNA.2PubMed Central. Saliva DNA quality and genotyping efficiency in a predominantly elderly population So the wet-lab portion of the process, the actual reading of your genetic markers, is not where the two companies meaningfully differ.

The real divergence happens downstream, when algorithms assign your DNA segments to population groups or estimate disease risk. That is where the reference databases, the statistical assumptions, and the company’s particular scientific choices all come into play.

Ethnicity Estimates and Why They Disagree

The most visible product of a consumer DNA test is the pie chart or map showing your ancestry composition: 42% English, 18% West African, 12% Indigenous American, and so on. These estimates are the feature people most commonly compare between AncestryDNA and 23andMe, and they are also where the two services diverge the most.

A study that tested identical twins through multiple DTC companies illustrates this clearly. When twin pairs were both tested by the same company, the mean agreement in ancestry percentages ranged from about 94.5% to 99.2%, which is what you’d expect since identical twins share nearly all their DNA. But when each twin was tested by a different company, mean agreement dropped to between roughly 53% and 84%.3PubMed Central. Consistency of Direct to Consumer Genetic Testing Results Among Identical Twins The same genome, filtered through different algorithms and reference panels, produced substantially different ancestry breakdowns.

This does not mean one company got it right and the other got it wrong. Both are making probabilistic estimates based on comparing your DNA to their curated reference populations. AncestryDNA and 23andMe each built their reference panels from different sets of people with documented multi-generational ties to specific regions. If one company has more reference samples from, say, Scandinavia, it can distinguish between Norwegian and Swedish ancestry more finely, while the other might lump them together as “Northwestern European.” Neither answer is wrong; they just reflect different granularity.

Why Your Results Change When Companies Update

If you’ve had your DNA tested for a few years, you may have noticed your ethnicity estimates shift after an update. AncestryDNA has rolled out several major updates to its ethnicity algorithm, and 23andMe has done the same. A region that once showed as 15% of your ancestry might drop to 8% or disappear entirely, while a new region you’d never seen before suddenly appears at 10%.

This happens because the companies periodically expand and refine their reference panels. When they add more reference samples from a previously underrepresented region, the algorithm gets better at distinguishing that population’s genetic signature from nearby groups. The practical effect is that your raw DNA hasn’t changed, but the yardstick the company measures it against has. A segment that was previously labeled “broadly Southern European” might get reclassified as “Sardinian” once the company adds enough Sardinian reference genomes to tell the difference.

The quality of genotype imputation, the process of statistically inferring genetic variants that weren’t directly measured on the chip, also depends on the reference panel. Research has shown that imputation accuracy varies substantially depending on which reference panel is used and how genetically similar the target population is to that panel. For example, a study in an underrepresented Southeast Asian population found that imputation accuracy and the number of high-confidence variant sites differed considerably across major reference panels like TOPMed, 1000 Genomes, and others.4Scientific Reports. A diverse ancestrally-matched reference panel increases genotype imputation accuracy in a underrepresented population The choice of reference data isn’t a minor detail; it fundamentally shapes the results.

Accuracy Gaps for Non-European Ancestry

Both AncestryDNA and 23andMe perform best for people with predominantly European ancestry, and this is one of the most important caveats that the marketing materials understate. The genetic research enterprise, including the genome-wide association studies that underpin both ancestry estimation and health risk scoring, has historically drawn most of its participants from European-descended populations. That imbalance trickles down to every consumer product built on that data.

For ancestry estimation specifically, research on local ancestry inference in admixed populations has found that accuracy varies by the ancestral component being identified. In populations with mixed European, African, and Indigenous American ancestry, European tracts were correctly identified about 96 to 99% of the time, African tracts at 98 to 99%, but Indigenous American tracts only at about 88 to 94%. When the algorithm made errors on Indigenous American segments, it most frequently misassigned them as European.5PubMed Central. Characterizing features affecting local ancestry inference performance in admixed populations For someone with significant Indigenous American heritage, this means both AncestryDNA and 23andMe are likely to undercount that component and overcount European ancestry.

This isn’t a flaw unique to either company. It reflects the state of the underlying science. Populations that have been studied more intensively have richer reference data, which means algorithms can distinguish finer subgroups and assign segments more confidently. Populations with less representation in reference panels get broader, less precise labels, and some of their ancestry gets absorbed into better-characterized neighboring groups.

AncestryDNA has invested in expanding its reference panel for African, Asian, and Indigenous populations, and 23andMe has similarly tried to diversify. But neither can fully compensate for decades of research bias in genomics. If your background includes ancestry from regions historically underrepresented in genetic studies, you should expect more uncertainty in your results regardless of which company you choose.

Health Reports and the False-Positive Problem

23andMe is the more prominent player in health-related genetic testing, offering FDA-authorized reports on carrier status for conditions like cystic fibrosis and sickle cell disease, as well as genetic health risk reports for conditions including late-onset Alzheimer’s disease and Parkinson’s. AncestryDNA has offered health features at various points but has historically focused more on genealogy and ethnicity. When comparing health accuracy, the conversation usually centers on 23andMe and similar DTC health offerings.

A study that examined variants reported in DTC raw data found that 40% of the variants analyzed turned out to be false positives when checked with clinical-grade confirmation testing. Of those false positives, about 94% were in cancer-related genes.6Genetics in Medicine. False-positive results released by direct-to-consumer genetic tests highlight the importance of clinical confirmation testing for appropriate patient care Some well-known pathogenic variants, like the Ashkenazi Jewish BRCA1/2 founder mutations, were confirmed in every case. But other variants, including several in the tumor-suppressor gene TP53, were not confirmed at all.

This finding deserves some context. The 40% false-positive rate applied to raw data variants that consumers or third-party tools pulled from the unfiltered genotyping output, not necessarily to the curated, FDA-reviewed reports that 23andMe formally presents to users. The FDA authorization process for specific health reports involves a higher standard of validation. But many users download their raw data and upload it to third-party interpretation services, and that is where the false-positive risk becomes a real clinical concern. If you find a scary-looking variant in your raw data, getting clinical confirmation testing before making any medical decisions is essential.

Polygenic Risk Scores and the Ancestry Gap

Beyond single-gene variants, both companies and many third-party tools are increasingly interested in polygenic risk scores, which combine the small effects of hundreds or thousands of genetic variants to estimate your risk for common diseases like heart disease, type 2 diabetes, or breast cancer. These scores are where the ancestry accuracy gap becomes most consequential for health.

Polygenic risk scores work best when the person being scored is genetically similar to the population that was used to develop the score. Research using a large, diverse biobank found that score accuracy decreases on an individual-by-individual basis as genetic distance from the training population increases. Among people with European ancestry, those in the most genetically distant group had about 14% lower accuracy compared with the closest group. People with Hispanic or Latino American ancestry who were genetically closest to the European training set performed comparably to the most distant Europeans, meaning the drop-off for more genetically distant individuals in non-European groups was even steeper.7Nature. Polygenic scoring accuracy varies across the genetic ancestry continuum

Multiple reviews have confirmed this pattern. Polygenic risk scores have lower accuracy when applied across populations that are genetically distant from the original discovery sample, regardless of the method used to build the score.8PubMed Central. Principles and methods for transferring polygenic risk scores across global populations The reasons include differences in which genetic variants are common in different populations, differences in the patterns of neighboring variants that travel together on chromosomes, and differences in environmental factors that interact with genetic risk.9PubMed Central. Polygenic risk score translation across diverse populations

For a consumer, this means that if you have non-European ancestry, polygenic risk scores from either 23andMe or any third-party tool analyzing AncestryDNA data are likely to be less reliable for you. This isn’t something either company can solve alone; it reflects a structural problem in genomic research that the field is slowly working to address by conducting larger studies in diverse populations.

Relative Matching and Database Size

For many users, the most practically valuable feature of a DNA test is not the ethnicity pie chart but the list of genetic relatives, people who share enough DNA with you to be identified as cousins. Here, the key differentiator between AncestryDNA and 23andMe is not accuracy in any technical sense but database size.

AncestryDNA has the larger database, with over 25 million people tested as of recent estimates. 23andMe’s database is also substantial but smaller. The practical implication is straightforward: the more people who have tested with a given company, the more potential matches you’ll find there.10The DNA Geek. Estimating the Sizes of the Genealogical atDNA Databases For adoptees searching for biological family, for genealogists trying to break through brick walls, or for anyone curious about distant cousins, database size matters more than any algorithmic nuance.

AncestryDNA also integrates its DNA matches with the company’s massive collection of historical records, family trees, and other genealogical data, which gives it a significant edge for genealogical research specifically. 23andMe, by contrast, offers more robust health reporting alongside its relative matching. The choice between them often comes down to what you want to do with the results rather than which one reads your DNA more accurately.

Regulation and Oversight

The regulatory landscape for DTC genetic testing is still catching up with the industry. In the United States, the FDA has authorized specific health-related tests from 23andMe, meaning those particular reports have been reviewed for analytical and clinical validity. But the broader category of DTC genetic testing, including ancestry estimation and raw data access, operates with relatively light regulatory oversight. A scoping review of DTC genetic testing regulations noted the need for policies requiring clinical validity assessments before tests reach the public, controls on marketing claims, and requirements for healthcare provider involvement.11PubMed Central. Considerations for developing regulations for direct-to-consumer genetic testing: a scoping review using the 3-I framework

For ancestry results specifically, there is no regulatory body that checks whether a company’s ethnicity estimates are “correct.” The percentages you see are model outputs, not medical diagnoses, and different models will produce different results. No authority certifies that any company’s reference panel is adequate or that its algorithm meets a specific accuracy threshold for ancestry estimation.

Privacy and Law Enforcement Access

Accuracy isn’t the only consideration when choosing between these services. Both companies store your genetic data, and the policies governing who can access it differ. 23andMe has faced financial difficulties in recent years, raising questions about what happens to its database if the company is sold. AncestryDNA, as part of a larger genealogy company, has its own data-sharing considerations.

A broader concern applies to any consumer DNA database. Law enforcement has used genetic genealogy techniques to identify criminal suspects by matching crime-scene DNA against databases of consumers’ distant relatives. By finding a partial match with someone’s third or fourth cousin, investigators can narrow suspects down to a family and eventually an individual.12PubMed Central. Commercial DNA tests and police investigations: a broad bioethical perspective Both AncestryDNA and 23andMe have stated that they require valid legal process before sharing data with law enforcement, but the risk that sensitive information surfaces is greater in this investigative genetic genealogy context than in traditional forensic databases.

When you submit your DNA, you’re also generating information about your biological relatives, people who never consented to having their genetic data inferred. This is true regardless of which company you choose, and it’s worth thinking about before you spit in the tube. The accuracy question may be what draws people in, but the privacy question is what keeps ethicists up at night.