What Is Variant Annotation and Why Is It Important?

Variant annotation is the process of attaching biological and clinical meaning to the millions of DNA differences found in a person’s genome. When a genome is sequenced, the raw output is a long list of positions where an individual’s DNA differs from a reference. On its own, that list is almost useless. Variant annotation layers on information about where each change falls in the genome, what effect it could have on a protein or regulatory element, how common it is in human populations, and whether it has been linked to disease. Without this step, genomic sequencing would produce data but no answers.

From Raw Variants to Usable Information

A typical human genome contains around four to five million variants compared to the reference sequence. The vast majority are harmless. Variant annotation is essentially a triage system that helps researchers and clinicians sort through that haystack to find the variants that actually matter. The process typically involves running specialized software that cross-references each variant against dozens of databases, prediction algorithms, and classification criteria. Modern annotation pipelines can attach hundreds of data points to a single variant. One recently described pipeline, for instance, consolidates output from multiple annotation engines and plugins into files with up to 920 columns of information per variant, covering everything from predicted protein damage to splice-site disruption to clinical classification.

1PubMed Central. MuSA: a Nextflow pipeline for deep, reproducible annotation and clinical ranking of genomic variants

The first and most basic layer of annotation identifies where a variant sits relative to known genes. A change that falls inside a gene’s protein-coding region gets flagged differently from one that sits between genes. Within coding regions, the annotation distinguishes between changes that swap one amino acid for another (missense variants), changes that create a premature stop signal (nonsense variants), and changes that shift the reading frame of the protein (frameshift variants). Each of these categories carries different implications for protein function.

Beyond coding regions, annotation tools also flag variants that land in splice sites, which are the boundaries where segments of a gene are stitched together. A variant that disrupts a splice site can cause an entire chunk of protein to be skipped or a nonsensical chunk to be included. Variants in regulatory regions, promoters, and enhancers are harder to interpret but increasingly recognized as important, especially as large-scale experiments map which stretches of non-coding DNA actually control gene activity.

Computational Prediction Scores

Once a variant is categorized by location, the next question is whether it actually does damage. This is where pathogenicity prediction tools come in. These algorithms estimate the likelihood that a given variant disrupts normal protein function or gene regulation. They draw on information like evolutionary conservation (has this position stayed the same across species for millions of years?), the physical and chemical properties of the amino acids involved, and protein structural context. A protein performs its function by folding into a precise three-dimensional shape held together by a web of chemical interactions between amino acid residues, and a single substitution in the wrong spot can destabilize the whole structure.

2Briefings in Bioinformatics. Interpreting functional effects of coding variants: challenges in proteome-scale prediction, annotation and assessment

Not all prediction tools perform equally. A study comparing tools on somatic (cancer-related) variants found that CADD, REVEL, PolyPhen-2, and PROVEAN were among the top performers, while SIFT, one of the most commonly used tools, fell into a second tier. Combining tools two by two only marginally improved accuracy, largely because the tools often disagreed with each other.

3PubMed. Comparison of Pathogenicity Prediction Tools on Somatic Variants

A separate evaluation using clinically relevant variant datasets reinforced that meta-predictors, which combine signals from multiple underlying tools, tend to outperform individual ones. REVEL and ClinPred both showed strong discriminatory power. However, there was a catch: while these tools performed impressively on well-characterized benchmark datasets, their specificity dropped considerably when applied to variants encountered in real clinical settings. REVEL’s specificity, for example, fell from 0.95 on a standard benchmark to 0.60 on a clinical dataset. That gap matters because lower specificity means more benign variants get incorrectly flagged as harmful, potentially leading to unnecessary clinical concern.

4PubMed Central. Assessing performance of pathogenicity predictors using clinically relevant variant datasets

The takeaway for anyone relying on these scores is that no single tool should be treated as definitive. Prediction scores are one input into variant interpretation, not the final word. Clinical laboratories routinely use multiple tools and weigh their outputs alongside other evidence.

Population Frequency Databases

One of the most powerful filters in variant annotation is simple: how common is this variant in the general population? If a variant shows up frequently in healthy people, it is overwhelmingly likely to be benign. This logic drives the use of population databases like the Genome Aggregation Database (gnomAD), which aggregates sequencing data from hundreds of thousands of individuals across multiple ancestry groups.

5PubMed. High prevalence of cancer-associated TP53 variants in the gnomAD database: A word of caution concerning the use of variant filtering

The recommended approach for filtering out common variants is to use the “population maximum” allele frequency, which looks at the highest frequency of a variant across major continental ancestry groups. If a variant is common in any one population, the reasoning goes, it can generally be assumed benign across all populations.

6PubMed Central. Variant interpretation using population databases: Lessons from gnomAD

Frequency-based filtering sounds straightforward, but it depends on how accurately those frequencies are estimated, and that accuracy depends on who was included in the database. Recent work using local ancestry inference in gnomAD revealed that for over 80% of genomic sites examined, the refined ancestry-based allele frequencies were higher than previously estimated population-group maximums. Among ClinVar variants where the new frequency exceeded the old estimate, roughly 88 to 89% were classified as benign or likely benign. In other words, better ancestry-aware frequency calculations can help reclassify variants more accurately and reduce false alarms.

7Nature Communications. Improved allele frequencies in gnomAD through local ancestry inference

ClinVar and the Value of Shared Interpretations

While population databases tell you how common a variant is, ClinVar tells you what other laboratories and researchers have concluded about it. ClinVar is a free, public archive that collects submissions from clinical labs, research groups, and expert panels, each reporting their interpretation of a variant’s relationship to disease. For any given variant, ClinVar aggregates these submissions and flags whether the interpretations agree or conflict.

8PubMed Central. Using ClinVar as a Resource to Support Variant Interpretation

The value of ClinVar is cumulative. A variant submitted by one lab with uncertain significance may be reclassified once five other labs contribute additional evidence. But the system also surfaces disagreements. When one lab calls a variant pathogenic and another calls it benign, ClinVar makes that conflict visible, which is itself a form of useful annotation. It signals that the evidence is contested and that additional scrutiny is warranted before acting on the result clinically.

The Five-Tier Classification System

The culmination of variant annotation in a clinical setting is formal classification. The widely adopted framework, developed jointly by the American College of Medical Genetics and Genomics (ACMG) and the Association for Molecular Pathology (AMP), sorts variants into five categories: pathogenic, likely pathogenic, uncertain significance, likely benign, and benign. Classification draws on multiple types of evidence, including population data, computational predictions, functional laboratory studies, and segregation patterns in families.

9PubMed Central. Standards and Guidelines for the Interpretation of Sequence Variants: A Joint Consensus Recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology

In practice, though, different laboratories applying these same guidelines can reach different conclusions. A study that had three independent labs classify 158 variants found complete five-category agreement for only about 54% of them. For roughly 11%, the disagreement was clinically significant, meaning one lab called a variant pathogenic or likely pathogenic while another called it a variant of uncertain significance or benign. That kind of discrepancy could mean the difference between a patient receiving a diagnosis and treatment versus being told nothing actionable was found.

10The American Journal of Human Genetics. Variant Classification Concordance using the ACMG-AMP Variant Interpretation Guidelines across Nine Genomic Implementation Research Studies

The “variant of uncertain significance” category, often abbreviated VUS, is the bane of clinical genomics. It means the evidence is insufficient to call a variant harmful or harmless. Patients who receive a VUS result are often left in limbo, unsure what it means for their health. Improving annotation, whether through better prediction tools, larger population databases, or more functional data, is one of the main ways the field works to shrink the VUS pile over time.

Variant Annotation in Cancer

Cancer genomics uses variant annotation differently from rare-disease genetics. Tumor sequencing identifies somatic variants, changes that arose in the cancer cells rather than being inherited. The clinical question shifts from “does this variant cause a genetic disease?” to “does this variant drive the tumor, and is there a drug that targets it?” Separate guidelines from AMP, the American Society of Clinical Oncology, and the College of American Pathologists provide a tiered evidence system for classifying somatic variants by their clinical actionability.

11PubMed Central. Variant Interpretation for Cancer (VIC): a computational tool for assessing clinical impacts of somatic variants

Knowledge bases like OncoKB annotate somatic mutations with their biological effect, prognostic significance, and whether any approved or investigational therapies target them. Treatment implications are stratified by the strength of evidence, from FDA-approved indications at the top to preliminary research findings at the bottom.

12PubMed. OncoKB: A Precision Oncology Knowledge Base

This is where variant annotation translates most directly into treatment decisions. If a lung tumor carries a specific mutation in the EGFR gene, for instance, annotation with OncoKB can point the oncologist toward a targeted therapy with strong evidence of benefit. Without that annotation layer, the raw mutation data would sit in a file, unconnected to clinical action.

Pharmacogenomics and Drug Response

A related but distinct application of variant annotation is pharmacogenomics, which links inherited genetic variants to differences in drug metabolism and response. The Pharmacogenomics Knowledgebase (PharmGKB) curates evidence about which genetic variants affect how patients respond to specific medications, assigning levels of evidence to each gene-drug association. It provides dosing guidelines and annotated drug labels that clinicians can use when a patient’s genotype is known.

13PubMed Central. Pharmacogenomics knowledge for personalized medicine

A patient who carries certain variants in drug-metabolizing genes might break down a medication too quickly (rendering it ineffective) or too slowly (causing toxic buildup). Annotating these variants and linking them to PharmGKB’s curated evidence allows prescribers to adjust doses or choose alternative drugs before a patient experiences an adverse reaction. Evidence from the pharmacogenomic literature is curated into PharmGKB as variant annotations, which feed into clinical recommendations.

14PubMed Central. An Evidence-Based Framework for Evaluating Pharmacogenomics Knowledge for Personalized Medicine

Structural Variants Are Harder to Annotate

Most annotation tools and databases are optimized for single-nucleotide changes and small insertions or deletions. Structural variants, which involve larger rearrangements of DNA such as deletions, duplications, inversions, or translocations affecting hundreds to millions of base pairs, present a different challenge. They can disrupt multiple genes at once, alter regulatory landscapes, or create entirely new gene fusions.

A systematic benchmarking of eight structural variant prioritization tools found that while both knowledge-driven and data-driven approaches showed comparable effectiveness in predicting pathogenicity, performance varied substantially across tools and genomic contexts. The study emphasized that choosing the right tool depends on the specific research question, and no single tool dominated across all scenarios.

15PubMed Central. Systematic assessment of structural variant annotation tools for genomic interpretation

The Reference Genome Problem

A subtle but consequential issue in variant annotation is which version of the human reference genome is used. The two versions in widest use, GRCh37 (also known as hg19) and GRCh38 (hg38), are largely identical but differ in complex regions that were revised and updated in the newer build. When the same sequencing data is analyzed against both references, a study found that about 1.5% of single-nucleotide variants and 2% of insertions or deletions were discordant, meaning they were detected on one reference but not the other. Most of these discordant calls were concentrated in discrete genomic windows rather than spread randomly across the genome.

16American Journal of Human Genetics. Choice of reference genome assembly significantly impacts variant calling and interpretation

When laboratories need to convert variant coordinates from one reference build to another, they use “liftover” tools. An evaluation of three widely used tools found that all three successfully converted more than 99% of ClinVar variants from GRCh37 to GRCh38. The small number of variants that failed conversion across all three tools were all insertions, deletions, or duplications, and nearly all were classified as benign or likely benign, with only one pathogenic or likely pathogenic variant among them.

17PubMed Central. Evaluation of Liftover Tools for the Conversion of Genome Reference Consortium Human Build 37 to Build 38 Using ClinVar Variants

The practical implication is that reference genome choice rarely causes dramatic problems for well-studied coding variants, but it can matter in edge cases, particularly for structural variants and variants in recently revised regions. Labs that have not yet migrated to GRCh38 may be missing a small fraction of true variants or calling false ones in specific genomic windows.

Ancestry Gaps and the Risk of Misdiagnosis

One of the most consequential limitations of variant annotation has been its dependence on databases that historically over-represented people of European ancestry. When a variant that is common in one ancestry group but absent from the reference databases shows up in a patient’s results, it may be misclassified as rare and potentially pathogenic simply because the database never included enough people from that background to recognize it as benign.

A landmark study documented exactly this problem with cardiac genetic testing. Multiple patients of African ancestry received reports calling certain variants pathogenic. Those variants were later reclassified as benign once population data from Black Americans was incorporated. The study found that the misclassified variants were significantly more common among Black Americans than white Americans, and simulations showed that including even small numbers of Black Americans in the original reference databases probably would have prevented the misclassifications entirely.

18PubMed Central. Genetic Misdiagnoses and the Potential for Health Disparities

This is not a historical curiosity. Although databases like gnomAD have grown substantially more diverse, gaps remain, particularly for populations in Africa, South Asia, the Middle East, and indigenous communities worldwide. Every variant annotation performed today is only as reliable as the reference data behind it. For patients from underrepresented groups, the rate of variants of uncertain significance tends to be higher, and the risk of misclassification does not fully go away until those databases catch up.

Deep Learning and the Shift Toward Empirical Data

The field of variant annotation is in the middle of a broad shift. Earlier tools relied heavily on theoretical predictions, estimating damage based on evolutionary conservation and amino acid chemistry. Increasingly, new tools incorporate empirically generated data, including protein structure predictions and large-scale functional experiments, rather than relying solely on theoretical models.

AlphaMissense, an adaptation of the protein-structure prediction system AlphaFold, represents one of the most prominent examples. It was fine-tuned on human and primate variant frequency data to predict missense variant pathogenicity, combining structural context and evolutionary conservation. It achieved strong results across genetic and experimental benchmarks without being explicitly trained on disease-labeled data.

19PubMed. Accurate proteome-wide missense variant effect prediction with AlphaMissense

Experimental approaches are also advancing. Saturation prime editing, a technique that uses CRISPR-based editing to systematically introduce every possible single-nucleotide change across a gene, can generate functional classifications for thousands of variants at once. One demonstration applied this to the NPC1 gene, mutations in which cause Niemann-Pick disease type C, creating a comprehensive functional map of the gene’s variants.

20PubMed Central. Saturation variant interpretation using CRISPR prime editing

These experimental datasets are valuable precisely because they do not depend on whether a variant has been observed in a patient or a population database. They provide functional evidence from scratch, which can resolve variants of uncertain significance that prediction tools and population frequency alone cannot settle.

Rare Disease Diagnosis at Scale

For families seeking answers about a rare genetic condition, variant annotation is the engine that makes diagnosis possible. The process typically begins with sequencing (often of the affected individual and both parents), followed by annotation and filtering to narrow millions of variants down to a handful of candidates. Platforms designed for this purpose combine variant annotation with tools for filtering by inheritance pattern, gene-disease association, and predicted functional impact.

One such platform, seqr, has been used to analyze over 10,000 families at a single center, supporting the diagnosis of more than 3,800 individuals with rare diseases and contributing to the discovery of over 300 novel disease genes.

21PubMed Central. seqr: A web-based analysis and collaboration tool for rare disease genomics

Those numbers illustrate both the power and the current ceiling of the approach. Even with sophisticated annotation, fewer than 40% of the families analyzed at that center received a definitive molecular diagnosis. The remaining 60% had variants that could not be confidently linked to disease, often because the gene had not yet been associated with any known condition, or because the variant’s effect was genuinely uncertain. Improving annotation is one of the primary ways to close that diagnostic gap.

Applications in Agriculture and Livestock

Variant annotation is not limited to human medicine. In agricultural genomics, annotating genetic variants against regulatory and functional maps of the genome helps identify the specific DNA changes underlying traits like disease resistance, growth rate, and yield. In livestock, integrating variant annotations with epigenomic data has been shown to identify the tissues and cell types underlying economically important traits and to improve the accuracy of genomic prediction models used in breeding programs.

22PubMed. Annotation and assessment of functional variants in livestock through epigenomic data

In plant genomics, a similar logic applies. Rare-allele variants in crops like soybean are of particular interest because they may be linked to stress resistance, yield, and environmental adaptability, traits that breeders want to select for but that standard breeding approaches may overlook if the underlying variants are too uncommon to appear in small panels.

23PubMed Central. Landscape of rare-allele variants in cultivated and wild soybean genomes

The core logic is the same across species: sequencing generates a list of variants, and annotation transforms that list into biological and practical meaning. What changes between human clinical work and agricultural breeding is the set of databases, the downstream decisions, and the regulatory framework, but the annotation step itself remains the critical bridge between raw sequence data and actionable knowledge.