How to Read Mutation Codes and What They Mean

Mutation codes follow a standardized naming system that tells you exactly where in a gene a change occurred and what that change looks like at the molecular level. The system, known as HGVS nomenclature, has become the international standard for describing genetic variants, and once you understand a few basic conventions, even intimidating strings like “c.1521_1523delCTT” or “p.Val600Glu” become readable. The logic is consistent: a prefix tells you which molecule you’re looking at, a number tells you the position, and a short code tells you what changed.

Why There Is a Standard System at All

For years, researchers described genetic changes in whatever format they liked. A variant discovered in one lab might be written one way in their paper and a completely different way in another group’s publication, even when both teams were describing the same change. That inconsistency made it genuinely hard to compare findings or translate research into clinical care.1PubMed Central. Standard mutation nomenclature in molecular diagnostics: practical and educational challenges The Human Genome Variation Society proposed a unified naming system around 2000, and it has since been adopted globally as the default way to write variant descriptions.2PubMed. HGVS Recommendations for the Description of Sequence Variants: 2016 Update The rules are maintained by a committee under HUGO, the Human Genome Organisation, and they get updated periodically through a community consultation process.3PubMed Central. HGVS Nomenclature 2024: improvements to community engagement, usability, and computability

You’ll see the system called “HGVS nomenclature” in lab reports, scientific papers, and genetic databases. If you encounter a mutation code on a test result or in a journal article written in the last decade or so, it almost certainly follows these conventions.

The Prefix Tells You Which Molecule

Every HGVS variant description starts with a short prefix followed by a period. That prefix is the single most important piece to read first, because it tells you whether the change is being described at the level of DNA, RNA, or protein. Here are the ones you’ll encounter most often:

  • c. Coding DNA. The position numbers refer to the gene’s coding sequence, starting from the first nucleotide of the start codon. This is the most common prefix on clinical genetic test reports.
  • g. Genomic DNA. Positions refer to a full genomic reference sequence, including non-coding regions. You’ll see this in research papers and large-scale genome analyses.
  • p. Protein. The description tells you what happened to the amino acid sequence. Protein-level codes use three-letter or one-letter amino acid abbreviations instead of DNA bases.
  • r. RNA. Describes changes at the RNA transcript level. Less commonly seen in routine clinical reports but important when splicing effects are being discussed.
  • m. Mitochondrial DNA. Same logic as genomic DNA, but specific to the mitochondrial genome.

When you see a variant written as “c.1234A>G,” you immediately know someone is talking about a change in the coding DNA sequence at position 1234. When you see “p.Glu600Val,” you know the description is at the protein level, telling you about an amino acid swap at position 600.

Reading DNA-Level Codes

DNA-level descriptions (those starting with c. or g.) come in a few standard patterns. The three most common types of changes are substitutions, deletions, and insertions.

Substitutions

A substitution is a single-letter swap. The format is: position, original base, a greater-than sign, and the new base. So “c.100A>G” means that at position 100 of the coding sequence, an adenine (A) was replaced by a guanine (G). The greater-than sign (>) always means “changed to.” If you see it in a code, you’re looking at a substitution.

Deletions

A deletion means one or more bases were lost. The format uses “del” after the position. “c.100delA” means the adenine at position 100 was deleted. If multiple consecutive bases are missing, you’ll see a range: “c.100_102del” means positions 100 through 102 are gone. Sometimes the deleted bases are spelled out after “del” (like “c.100_102delACT”), and sometimes they’re left off since the reference sequence already tells you what was there.

Insertions and Duplications

An insertion means extra bases were added between two existing positions. “c.100_101insGGG” means three guanines were inserted between positions 100 and 101. Notice that insertions always name two flanking positions separated by an underscore, because the new bases sit between existing ones.

A duplication is a special case: when the inserted sequence is an exact copy of what’s immediately upstream, the system uses “dup” instead of “ins.” So if a stretch of bases gets repeated, you’ll see something like “c.100_102dup” rather than an insertion notation. This distinction matters because duplications and insertions can arise through different biological mechanisms, and the “dup” label helps researchers and clinicians recognize the pattern.

Deletion-Insertions

Sometimes bases are deleted and replaced by different bases in a single event. The format combines both: “c.100_102delinsGGTT” means positions 100 through 102 were removed and replaced with GGTT. The old sequence and the new sequence don’t have to be the same length.

Reading Protein-Level Codes

Protein-level descriptions start with “p.” and use amino acid names instead of DNA bases. You’ll see amino acids written either in three-letter codes (like Gly, Ala, Val) or single-letter codes (G, A, V). Three-letter codes are standard in clinical reports because they’re less error-prone. The general format is: original amino acid, position number, new amino acid.

Missense Changes

A missense variant swaps one amino acid for another. “p.Val600Glu” means that at position 600 of the protein, valine was replaced by glutamic acid. This is one of the most commonly discussed types of variant because a single amino acid swap can dramatically change how a protein folds or functions. The well-known BRAF V600E mutation in cancer, for instance, follows this exact pattern.

Nonsense Changes

“p.Arg100Ter” (or “p.Arg100*”) means that at position 100, the arginine codon was replaced by a stop signal, abbreviated “Ter” or shown as an asterisk. The protein gets cut short at that point. These are called nonsense variants, and they tend to have severe effects because a truncated protein usually can’t do its job.

Frameshifts

When an insertion or deletion at the DNA level shifts the reading frame, the protein description uses “fs” to flag it. “p.Leu100fs” tells you that starting at leucine 100, the reading frame shifted and everything downstream produces a garbled amino acid sequence. Often you’ll also see “fsTer” followed by a number, like “p.Leu100fsTer15,” which means the shifted reading frame hits a premature stop signal 15 amino acids later.

Silent Changes

If a DNA change doesn’t alter the amino acid at all (because of redundancy in the genetic code), the protein-level description uses an equals sign: “p.Leu100=” means position 100 is still leucine despite a change in the DNA. These “synonymous” or “silent” variants were long assumed to be harmless, though research increasingly shows that some can affect how efficiently the protein is made.

Splice Site and Intronic Variants

Not all DNA changes sit neatly inside coding exons. Many clinically important variants occur at the borders between exons and introns, where the splicing machinery reads signals to cut and join the RNA transcript. HGVS notation handles these with an offset from the nearest exon boundary. “c.100+1G>A” means the change is one base into the intron, downstream of coding position 100. “c.200-2T>C” means it’s two bases into the intron, upstream of coding position 200.

The plus sign points into the intron on the downstream side of the exon, and the minus sign points into the intron on the upstream side. The positions closest to the exon (especially +1, +2, -1, and -2) are the most critical for splicing, and changes there frequently disrupt normal RNA processing. Research on whole-exome sequencing data suggests that positions like +2 and -2 may be even more important than +1 and -1 for splicing function, and that the splicing-relevant zone extends out to about the +9 and -9 positions.4PubMed Central. Intronic position +9 and -9 are potentially splicing sites boundary from intronic variants analysis of whole exome sequencing data Changes deeper into the intron are less likely to matter, though exceptions exist.

If you see a variant written as “c.100+5G>T,” you can read it as: five bases into the intron after coding position 100, a G was replaced by a T. Whether that actually disrupts splicing depends on the specific gene and context, but the notation itself always follows this offset logic.

What the Clinical Classification Means

Knowing how to read a mutation code is one thing. Knowing what it means for your health is another, and the code itself doesn’t tell you that. A variant’s clinical significance is assessed separately and assigned one of five standard categories, as recommended by professional genetics organizations:5PubMed Central. Standards and Guidelines for the Interpretation of Sequence Variants: A Joint Consensus Recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology

  • Pathogenic: The variant is considered disease-causing based on strong evidence.
  • Likely pathogenic: There is over 90% certainty the variant causes disease, but the evidence isn’t quite as airtight.
  • Uncertain significance (VUS): Not enough evidence exists to say whether the variant is harmful or harmless. This is the most frustrating result for patients and clinicians alike.
  • Likely benign: Over 90% certainty the variant does not cause disease.
  • Benign: Strong evidence that the variant is harmless.

The “likely” categories use a working threshold of greater than 90% certainty, which sounds high but still means roughly one in ten of those classifications could turn out to be wrong as more data comes in.5PubMed Central. Standards and Guidelines for the Interpretation of Sequence Variants: A Joint Consensus Recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology Variants of uncertain significance are especially common in genetic testing. If you receive a VUS result, it usually means more time or more data is needed before anyone can say what it does. Some VUS classifications get reclassified as pathogenic or benign years later as researchers study more families or run more functional experiments.

The classification system draws on multiple types of evidence: how common the variant is in the general population (rare variants are more suspicious), whether computer models predict it will damage the protein, whether lab experiments show a functional effect, and whether the variant tracks with disease in families. No single line of evidence is usually enough on its own.

Where to Look Up a Variant

If you have a variant code and want to know what’s been reported about it, the first stop is usually ClinVar, a freely accessible database run by the National Center for Biotechnology Information. ClinVar collects variant interpretations submitted by clinical labs and research groups. Each entry includes the HGVS description, the gene involved, the submitting lab’s classification, and supporting evidence.6Nucleic Acids Research. ClinVar: improving access to variant interpretations and supporting evidence Because multiple labs sometimes submit interpretations for the same variant, you can see whether there’s consensus or disagreement about its significance.

Other useful resources include gene-specific databases (called locus-specific databases) that catalog variants for particular genes in greater detail, and annotation tools like ANNOVAR that researchers use to cross-reference a list of variants against population frequency data, predicted functional effects, and known disease associations.7PubMed Central. Genomic variant annotation and prioritization with ANNOVAR and wANNOVAR These tools are more technical, but they illustrate an important point: reading the code is just the first step. Interpreting its meaning requires layering on additional information about the gene, the protein, the variant’s frequency, and the clinical context.

Common Pitfalls When Reading Mutation Codes

Even among professionals, mutation nomenclature trips people up. Several patterns cause repeated confusion.

Mixing up coding and genomic coordinates is one of the biggest sources of error. The same physical DNA change can have completely different position numbers depending on whether it’s described at the coding level (c.) or the genomic level (g.). If a paper says “c.100A>G” and another paper says “g.53271A>G,” those might be the same variant, just written against different reference sequences. Always check which reference sequence is being used before comparing position numbers across sources.

Another stumbling block is legacy nomenclature that predates the current standard. Older papers and some clinical databases still use outdated formats. You might see protein variants written as “V600E” without the “p.” prefix, or DNA changes described relative to an older numbering system. Research papers that first report novel variants often do not use standard nomenclature, which has been a persistent source of confusion in the field.1PubMed Central. Standard mutation nomenclature in molecular diagnostics: practical and educational challenges If a variant code looks odd or doesn’t parse by the rules described here, check whether it’s from an older publication using a now-superseded format.

A third pitfall is assuming the protein effect from the DNA code without checking. A DNA substitution might look benign at the DNA level but create a splicing problem that leads to a truncated protein. Conversely, a change that looks dramatic at the DNA level might be completely synonymous at the protein level. The DNA code and the protein code tell you different things, and both matter.

The Language Keeps Evolving

If you find conflicting conventions across different sources, it may be because the nomenclature rules get revised. The system went through a major update in 2016 and another round of improvements in 2024, with changes aimed at making the notation more computationally friendly and easier for software tools to parse.3PubMed Central. HGVS Nomenclature 2024: improvements to community engagement, usability, and computability Earlier versions of the rules handled certain edge cases differently, so a variant described in a 2005 paper may use slightly different formatting than the same variant in a 2024 clinical report.

The updates typically don’t change the core logic. Substitutions still use “>”, deletions still use “del”, protein changes still list the original and new amino acids. What changes are the finer details: how to handle complex rearrangements, how to describe variants in non-coding regulatory regions, what to do when a single DNA change has multiple downstream effects. For a general reader, the fundamentals covered here will get you through the vast majority of variant codes you’ll encounter on a genetic test report or in a research summary.

When Methylation Enters the Picture

Standard mutation codes describe changes to the DNA or protein sequence itself. But there’s a growing category of genetic testing that reports on methylation patterns rather than sequence changes. Methylation doesn’t alter the letters of the DNA code; instead, it adds chemical tags that influence whether a gene is turned on or off. Abnormal methylation at certain genomic regions is associated with a group of conditions known as imprinting disorders, where it matters whether a gene was inherited from the mother or the father.

Reporting methylation findings has its own consistency problem. Different labs have tested different sites within the same genomic region and described their results using different names, making comparisons across studies difficult. A European network working on imprinting disorders has proposed a separate nomenclature specifically for naming the regions where methylation is measured and for reporting whether methylation levels are normal or abnormal.8PubMed Central. Recommendations for a nomenclature system for reporting methylation aberrations in imprinted domains This system is distinct from HGVS variant nomenclature, but if you receive genetic testing results that mention methylation at differentially methylated regions, it’s worth knowing that standardization efforts are underway in that space too.

Methylation reports won’t use the c. or p. prefixes you’ve learned for sequence variants. Instead, they typically name the genomic region being tested and report a methylation ratio or percentage. The interpretation framework differs as well: rather than classifying a variant as pathogenic or benign, methylation testing usually asks whether the pattern is consistent with a specific syndrome. These are complementary types of genetic information, and increasingly, comprehensive genetic workups include both sequence analysis and methylation testing.