How Does DNA Genealogy Work? The Science of Ancestry

DNA genealogy works by comparing specific markers scattered across your genome to those of other people, looking for shared segments that indicate common ancestors. When two people share long, identical stretches of DNA, it means they inherited those stretches from the same ancestor, and the length of those shared stretches roughly indicates how far back that ancestor lived. The science behind this is surprisingly layered, combining molecular biology, statistics, and old-fashioned family-tree research into a system that can connect you to relatives you never knew existed or trace your maternal line back tens of thousands of years.

What Happens When You Swab Your Cheek

Most consumer DNA tests do not sequence your entire genome. Instead, they use a genotyping chip that reads hundreds of thousands of specific positions in your DNA where people commonly differ from one another. These positions are called single nucleotide polymorphisms, or SNPs, which are essentially spots where one person might carry a C and another an A. A large survey of direct-to-consumer genetic data documented over 7,000 genomes made public between 2011 and 2020, produced by more than six different testing companies, giving a sense of the scale at which consumer genotyping has operated.1PubMed Central. A survey of direct-to-consumer genotype data, and quality control tool (GenomePrep) for research Each company designs its chip slightly differently, choosing which SNPs to include, but the core idea is the same: read enough variable positions to build a genetic fingerprint that can be compared against everyone else in the database.

The testing companies then run your results through two main pipelines. The first is ancestry composition, which compares your SNP pattern to reference panels of people with well-documented ancestry from specific regions. If a stretch of your DNA looks statistically most like samples from, say, West Africa or Scandinavia, the algorithm assigns that segment to that population. The second pipeline is relative matching, which searches the database for people who share unusually long identical-by-descent segments with you. Those shared segments are the real engine of DNA genealogy.

How Shared DNA Connects You to Relatives

The fundamental principle is straightforward: when two people share a recent common ancestor, they both inherited pieces of that ancestor’s DNA. The closer the relationship, the more DNA they share. Full siblings share roughly half their autosomal DNA, first cousins share around an eighth, and by the time you get to third or fourth cousins, the shared amount has dwindled to small fragments. SNP-based technologies and massively parallel sequencing have expanded the analysis of distant kinship through the detection of these identity-by-descent segments and shared autosomal DNA.2PubMed Central. Single Nucleotide Polymorphisms in Distant Kinship Inference and Forensic Genetic Genealogy

The software scans your genome and a potential match’s genome, looking for stretches where both of you carry the same alleles on the same chromosome, unbroken by recombination. The longer and more numerous those stretches, the closer your relationship. This is more reliable than it might sound because random chance alone would not produce long matching segments between two unrelated people. A shared 30-centimorgan segment almost certainly points to a real common ancestor within the past few generations.

Where it gets interesting is at the margins. By the time you reach fifth or sixth cousins, many pairs share no detectable DNA at all despite having a documented genealogical connection. You both descend from the same person, but the shuffling of DNA through the generations chopped up the relevant segments until nothing identifiable remained. This is why testing companies emphasize that a lack of a DNA match does not necessarily mean two people are unrelated.

Why Siblings Get Different Ancestry Results

One of the most common surprises people encounter is that siblings sometimes get noticeably different ancestry breakdowns. This is not a glitch. Each child inherits a random combination of their parents’ DNA, and the randomness comes from recombination, the process where chromosomes swap segments during the formation of eggs and sperm. Two siblings each get about half of each parent’s DNA, but not the same half.

Research into the mechanics of this variation has shown that the genomic proportion two relatives share identically by descent varies depending on the specific history of recombination and segregation events in their pedigree.3PubMed Central. Variation in Genetic Relatedness Is Determined by the Aggregate Recombination Process For siblings, the expected sharing is 50 percent, but the actual number bounces around that average. In practical terms, this means one sibling might inherit more of a grandparent’s West African DNA while the other inherits more of the same grandparent’s European DNA, even though both grandparents contributed equally to both children’s genomes in a genealogical sense. The ancestry pie chart looks different because the DNA pie chart is different.

This variability compounds with each generation. Grandparent-grandchild relatedness has even more wiggle room than parent-child, and by the time you reach cousins, the variation is wide enough that some second cousins might share about as much DNA as typical third cousins, while others share an amount that looks more like first cousins once removed. The algorithms do their best to estimate the relationship, but they are working with probability, not certainty.

Tracing Maternal and Paternal Lines

Autosomal DNA, the kind that gets reshuffled every generation, is the workhorse for finding recent relatives. But two other types of DNA follow much simpler inheritance paths and are used for deeper lineage tracing.

Mitochondrial DNA passes from mother to child with almost no change across generations. Because it does not recombine the way autosomal DNA does, it preserves a remarkably clean record of maternal ancestry stretching back thousands of years. Researchers classify mitochondrial sequences into haplogroups, which are branches on a tree representing major migration events in human history. A study exploring mitochondrial DNA ancestry in a mixed Brazilian population, for example, used mtDNA as a stable genetic marker to investigate maternal ancestry, identifying that the most prevalent haplogroup categories were Native American and African, reflecting centuries of admixture.4PubMed Central. Revisiting the African mtDNA landscape through complete mitochondrial genomes When a consumer test reports your maternal haplogroup, it is placing you on that same deep tree.

The Y chromosome works similarly for paternal lineage but is available only in people who carry a Y chromosome. It passes from father to son largely intact, accumulating slow mutations over the millennia. Y-chromosomal short tandem repeats and single nucleotide polymorphisms are used to characterize paternal lineages, and because of their strict paternal inheritance, they have become a powerful means to trace paternal lineage evolution across populations.5PubMed. Genetic origin and forensic analysis of han populations from four districts in Shanghai, China, based on Y-chromosome STR and SNP markers Your Y-chromosome haplogroup tells you which branch of the global paternal tree you sit on, linking you to migration routes your direct male-line ancestors followed.

A key limitation of both lineage markers is that they trace only two thin threads of your ancestry. Your mitochondrial DNA tells you about your mother’s mother’s mother’s mother, and so on, but says nothing about the other ancestral lines converging in your family tree. Ten generations back, you have roughly a thousand ancestors, and your mtDNA reflects exactly one of them. The same is true for the Y chromosome and paternal lineage. These tools are powerful for understanding specific deep-time migration patterns, but they are not a complete ancestry picture.

Deep Ancestry and Ancient Migrations

One of the most captivating features of DNA genealogy is its ability to connect you to events tens of thousands of years in the past. Mitochondrial haplogroup L3, for instance, gave rise to all mitochondrial DNA sequences found outside the African continent. Every non-African maternal lineage traces back to this single ancestral haplogroup, making it a critical marker for understanding the out-of-Africa expansion that populated the rest of the world.4PubMed Central. Revisiting the African mtDNA landscape through complete mitochondrial genomes When your test places your haplogroup on the global tree, it is essentially showing you which branch of that ancient migration your direct maternal or paternal line followed.

Even more striking is the detection of archaic human ancestry. Modern genomic methods can identify segments of DNA inherited from Neanderthals and Denisovans, ancient human relatives who interbred with the ancestors of modern non-African populations. Researchers have developed methods to disambiguate Denisovan from Neanderthal ancestry segments in present-day humans and have applied them to 257 high-coverage genomes from 120 diverse populations, including 20 Oceanians with particularly high Denisovan ancestry.6PubMed Central. The Combined Landscape of Denisovan and Neanderthal Ancestry in Present-Day Humans Most people of European or East Asian descent carry roughly one to two percent Neanderthal DNA, while some populations in Melanesia carry an additional few percent of Denisovan DNA. Consumer tests now report this archaic ancestry as a percentage, giving you a window into interbreeding events that happened over 40,000 years ago.

Ancient DNA recovered from archaeological remains has expanded this picture even further. Applications of ancient DNA extend beyond ancestry reconstruction and population genetics to include forensic identification, kinship analysis, and pathogen detection in historical remains.7PubMed Central. Recovery and analysis of ancient DNA: challenges, methods, and applications in forensic and archaeological science Scientists have used ancient DNA to reveal that the people who built Stonehenge were largely replaced by a wave of migration from the continent, that the first Americans descended from populations that crossed from northeastern Asia, and that the Black Death pathogen traveled specific trade routes. This kind of work feeds back into the reference panels that consumer tests use, refining the picture of which populations lived where and when.

Filling in the Gaps with Imputation

Consumer genotyping chips read only a fraction of the roughly three billion positions in your genome. To get a more complete picture, the testing companies use a statistical technique that predicts the values of positions that were not directly measured. This works because DNA is not randomly arranged. Nearby positions on a chromosome tend to be inherited together, and large reference databases have mapped which variants tend to appear alongside which others. By comparing the SNPs your chip did read against these known patterns, algorithms can fill in many of the blanks with high confidence.

Genotype imputation relies on the principle that shared DNA segments among individuals are conserved over generations due to strong linkage between nearby markers and low recombination rates in certain regions.8PubMed Central. SNP Genotype Imputation in Forensics—A Performance Study The technique uses phased reference panels, which are databases of densely genotyped individuals whose chromosomal arrangements are already known. By matching your chip data against these references, the software infers what most of the unread positions likely are. This dramatically increases the number of usable data points for both ancestry estimation and relative matching, making cheap genotyping chips far more informative than they would be on their own.

Imputation is not perfect, especially for populations that are underrepresented in the reference panels. If the reference database has tens of thousands of individuals of northern European descent but only a few hundred from, say, Southeast Asia or indigenous South America, the imputation accuracy drops for those populations. This is one reason ancestry estimates are often more granular for European ancestry and vaguer for other regions, a problem the field is gradually addressing as reference panels become more diverse.

Forensic Genetic Genealogy

The same logic that connects you to a fourth cousin on a consumer platform can also connect crime-scene DNA to a suspect’s relatives, and through them, to the suspect. Forensic genetic genealogy applies advanced sequencing technologies to forensic DNA evidence and then uses genetic genealogy methods and genealogical research to produce possible identities of unknown perpetrators of violent crimes or unidentified human remains.9PubMed Central. Bridging Disciplines to Form a New One: The Emergence of Forensic Genetic Genealogy

The process typically goes like this: investigators extract DNA from crime-scene evidence and generate a SNP profile comparable to what a consumer test would produce. They upload that profile to a database that allows law enforcement searches. If a match appears, say a predicted third cousin of the unknown person, genealogists build out the family tree of that match using public records, vital records, and other genealogical tools. They work backward through the tree until they identify individuals who could plausibly be the unknown person, then investigators use traditional methods like surveillance or voluntary DNA collection to narrow it down.

The technique has been described as rivaling the forensic impact of STR analysis, which was introduced four decades ago and revolutionized criminal identification.10PubMed. Heating up three cold cases in Norway using investigative genetic genealogy Y-chromosome markers also play a role in forensic work, particularly in cases where standard autosomal profiling is not informative, such as sexual assault cases where male and female DNA are mixed together. Y-STR haplotypes can isolate the male contributor’s paternal lineage even in the presence of overwhelming amounts of female DNA.11PubMed Central. Forensic use of Y-chromosome DNA: a general overview

Privacy and the Reach of DNA Databases

The power of forensic genetic genealogy raises an uncomfortable question: how private is your DNA if a relative has uploaded theirs? A study leveraging genomic data from 600,000 individuals tested with consumer genomics found that roughly half of searches for European-descent individuals would turn up a third cousin or closer match, providing a search space small enough to permit re-identification using common demographic identifiers like age, location, and gender. The researchers projected that in the near future, virtually any European-descent person in the United States could be identifiable through this technique, whether or not they had ever taken a DNA test themselves.12bioRxiv. Re-identification of genomic data using long range familial searches

This creates a collective dimension to genetic privacy that is unlike most other personal data. Your decision not to test does not fully protect you if enough of your relatives have tested. The third cousin you have never met, living in another state, may have unknowingly placed a breadcrumb that leads back to your branch of the family. For forensic purposes, many people consider this an acceptable trade-off if it solves violent crimes. But the same technical capability could be used for other purposes entirely: insurance discrimination, paternity disputes, or state surveillance, depending on who controls the databases and what rules govern access.

Most consumer testing companies have policies limiting law enforcement access, and some databases were specifically designed to allow or disallow police searches based on user consent. But policies can change, companies can be acquired, and legal precedents around compelled database access are still being established. If you are weighing whether to test, the science itself is not the risk. The governance around the data is.

When DNA Results Can Be Misleading

DNA genealogy is remarkably reliable in most circumstances, but a few biological edge cases can produce confusing results. One is chimerism, a condition in which a single person carries two genetically distinct cell lines. This can happen naturally when fraternal twin embryos merge early in development, or it can result from medical procedures like bone marrow transplants. Transplant patients represent a particular challenge because their blood cells may carry the donor’s DNA rather than their own, meaning a blood-based DNA test could return a genetic profile that does not match the person who provided the sample.13Forensic Science International. Forensic implications of the presence of chimerism after hematopoietic stem cell transplantation

For consumer ancestry testing, chimerism is rare enough that most people will never encounter it. But for forensic applications, it matters. A crime-scene blood sample from a bone marrow recipient could point investigators toward the donor rather than the actual person. Awareness of this possibility has grown in forensic circles, and protocols now include checking for mixed profiles that might indicate chimerism.

Other limitations are more mundane but affect more people. Endogamy, where members of a community marry within the group for many generations, inflates the amount of shared DNA between individuals who are not closely related in the genealogical sense. In populations with long histories of endogamy, the algorithms can overestimate how closely two people are related because the shared segments are longer and more numerous than typical for the genealogical distance. Ashkenazi Jewish, Icelandic, and many island populations are well-known examples. If you come from such a community, your match list may be filled with predicted second and third cousins who are actually more distant.

Adoption, misattributed parentage, and donor conception can also produce results that do not align with a person’s understood family history. DNA does not lie about biological relationships, but it can reveal truths that family narratives have obscured. Testing companies have added warnings and support resources around these unexpected findings, recognizing that the emotional impact of a DNA match can be as significant as the scientific one.

What Ancestry Percentages Actually Represent

One of the most misunderstood features of consumer DNA tests is the ancestry composition pie chart. People tend to read “42 percent British and Irish” as a precise historical fact about their heritage, when it is really a statistical estimate based on how well stretches of their DNA match a particular reference panel. The reference panel itself is built from people who self-report that all four of their grandparents came from a specific region, which introduces its own assumptions about how neatly real populations map onto modern national borders.

These estimates improve over time as companies add more reference samples and refine their algorithms. If you tested in 2015 and again in 2024, your results may look different, not because your DNA changed, but because the reference data and the statistical model got better. A chunk of DNA that once could only be assigned to “Broadly European” might now be distinguishable as “French” versus “German” with more data to compare against. But the precision still has limits. Populations that lived close together and intermarried extensively, like neighboring countries in Western Europe, produce DNA patterns that are hard to tell apart. The test is more confident about broad continental ancestry than fine-grained regional ancestry for this reason.

Trace amounts, those small percentages below about five percent, deserve extra skepticism. Some of them are real signals of distant ancestry from unexpected places. Others are statistical noise produced by the algorithm’s difficulty in assigning short DNA segments to the correct reference panel. If your results show two percent of something surprising, it could be a genuine ancestor from that region several generations back, or it could be the algorithm hedging its bets on a few ambiguous segments. The companies themselves usually note confidence levels, and the higher-confidence estimates tend to collapse those trace amounts into broader categories.

How the Y Chromosome Has Reshaped Surname Research

Because the Y chromosome passes from father to son in much the same way that surnames do in many cultures, DNA genealogy has created an unexpected bridge between genetics and surname studies. Men who share a surname and a Y-chromosome haplogroup likely share a common male-line ancestor, even if no paper trail connects them. Y-STR profiling is widely used in this kind of paternal lineage tracing, since Y-chromosomal markers serve as direct records of the male line.5PubMed. Genetic origin and forensic analysis of han populations from four districts in Shanghai, China, based on Y-chromosome STR and SNP markers

Surname projects, organized by volunteers and hosted on genealogy platforms, collect Y-DNA results from men with the same or similar surnames and compare them. This has resolved longstanding questions about whether certain family branches are genuinely related or just happen to share a name. It has also revealed cases where a surname lineage was broken at some point by adoption, illegitimacy, or name changes, events invisible in paper records but obvious in the DNA. For people researching their family history, combining traditional genealogy with Y-DNA testing can break through brick walls that documents alone cannot.