AT and CG in DNA: What They Are and Why They Pair

Every strand of DNA is built from just four chemical letters, and they follow an inflexible pairing rule: adenine (A) always bonds with thymine (T), and cytosine (C) always bonds with guanine (G). This specificity comes from the shapes and hydrogen-bonding patterns of the four bases, which fit together like puzzle pieces that only have one correct match. The rule is simple to state, but its consequences ripple through everything from how accurately your cells copy their genome to why some organisms thrive in boiling hot springs.

What the Letters Actually Are

DNA’s four bases fall into two structural families. Adenine and guanine are purines, which have a two-ring molecular structure. Cytosine and thymine are pyrimidines, built on a single ring. When a base pair forms across the two strands of the double helix, it always consists of one purine paired with one pyrimidine. A purine paired with another purine would be too wide for the helix to accommodate uniformly, while two pyrimidines together would be too narrow, leaving a gap. The purine-pyrimidine combination keeps the helix at a consistent width, roughly the same distance between the two sugar-phosphate backbones at every rung of the ladder.

Researchers demonstrated this geometric constraint by building an unusual DNA-like molecule composed entirely of purines. In that system, the purine-purine pairs forced a wider distance between the backbone attachment points, and the molecule could only form a flat, ladder-type structure rather than a proper helix.1Cell Press (Chemistry & Biology). DNA Made of Purines Only The regular double helix you see in textbook illustrations depends on every rung being one big base plus one small base.

Why A Pairs With T and C Pairs With G

Beyond size, the bases are choosy about their partners because of hydrogen bonds. These are weak attractions that form when a hydrogen atom attached to a nitrogen or oxygen on one base lines up with a lone electron pair on the opposite base. The arrangement of donor and acceptor sites on adenine precisely matches those on thymine, forming two hydrogen bonds between them. Guanine and cytosine match up with three hydrogen bonds. Attempting to pair A with C, for example, puts donors opposite donors and acceptors opposite acceptors in the wrong positions, so stable bonds cannot form.

This selectivity is what makes DNA’s information storage reliable. When the cell copies a strand, the replication machinery reads each base on the template and slots in its complement on the new strand. Because the hydrogen-bond geometry only works for A-T and G-C, the system has a built-in error-checking mechanism at the chemical level.

Chargaff’s Rule and What It Tells You

Before anyone had worked out the double helix structure, the biochemist Erwin Chargaff noticed something striking. In 1950, he reported that in any sample of double-stranded DNA, the amount of adenine was essentially equal to the amount of thymine, and the amount of cytosine was essentially equal to the amount of guanine.2Oxford University Press. DNA sequence symmetries from randomness: the origin of the Chargaff’s second parity rule This observation, now called Chargaff’s first parity rule, was one of the key clues that led Watson and Crick to propose the base-pairing model in 1953.

Chargaff later extended the observation to individual strands: even on a single strand of DNA, the number of A’s roughly equals the number of T’s, and C’s roughly equal G’s. This second parity rule is less intuitive, since a single strand has no partner to enforce complementarity, and the reasons behind it are still debated. But the first rule is a direct, inevitable consequence of the pairing. If every A on one strand is matched by a T on the other, the totals have to be equal.

Two Hydrogen Bonds Versus Three

The fact that A-T pairs have two hydrogen bonds while G-C pairs have three has practical consequences you can measure in a lab. DNA with a higher proportion of G-C pairs is harder to pull apart because each rung of the ladder is held together by an extra bond. This shows up most clearly in thermal denaturation, the process of heating DNA until the two strands separate. The temperature at which half the DNA in a sample has separated (called the melting temperature) rises in a roughly linear fashion with increasing G-C content.3Nucleic Acids Research. Thermal stability of DNA

This relationship has biological implications. Prokaryotes that live in extreme heat were long predicted to have GC-rich genomes as an adaptation, and a large genomic analysis confirmed a positive correlation between GC content and growth temperature in these organisms.4PubMed Central. A positive correlation between GC content and growth temperature in prokaryotes The logic is straightforward: if your environment routinely approaches temperatures that would melt AT-rich DNA, packing in more GC pairs gives your genome a thermal buffer.

It Is Not Just About the Bonds Between Partners

A common simplification is that DNA’s stability comes entirely from hydrogen bonds across the two strands. In reality, a second force matters at least as much: base stacking. The flat, ring-shaped bases sit on top of one another along the helix like a stack of coins, and the interactions between adjacent bases in the same strand contribute substantially to holding the structure together.

How much each force contributes has been the subject of real scientific disagreement. One influential study concluded that stacking interactions are the dominant stabilizing force and that base pairing itself is actually slightly destabilizing for A-T pairs and contributes almost nothing for G-C pairs.5Nucleic Acids Research. Base-stacking and base-pairing contributions into thermal stability of the DNA double helix A later reanalysis using a different thermodynamic model reached the opposite conclusion, arguing that pairing contributions drive double-strand formation and that each hydrogen bond contributes roughly −0.72 kcal per mol of free energy.6bioRxiv. Base pairing and stacking contributions to double stranded DNA formation The debate reflects genuine complexity: separating stacking and pairing experimentally is extremely difficult because both happen simultaneously in any real DNA molecule.

For practical purposes, both forces matter, and neither alone explains why the double helix holds together. The hydrogen bonds between partners give DNA its information content by enforcing specific pairing. The stacking between neighbors gives the helix much of its mechanical rigidity.

Why Sequence Context Changes Everything

If A-T and G-C were the whole story, you could predict how stable a stretch of DNA is just by counting the ratio of the two pair types. But the identity of neighboring pairs matters too. Researchers have measured the stability of all ten possible combinations of two adjacent Watson-Crick pairs (called nearest-neighbor parameters), and the differences are large. At body temperature, the most stable two-step combination (GC next to itself) is substantially more stable than the least stable (TA next to itself).7PubMed. Improved nearest-neighbor parameters for predicting DNA duplex stability These parameters are the workhorse behind the software tools that researchers use to design primers, probes, and other short DNA sequences for laboratory work.

The nearest-neighbor model works well enough in standard lab buffers, but the inside of a cell is far from a clean salt solution. It is packed with proteins, metabolites, and other large molecules that create crowding effects. Updated nearest-neighbor parameters have been developed for these crowded conditions, and they produce more accurate predictions of how DNA actually behaves inside a living cell.8PubMed Central. Nearest-neighbor parameters for predicting DNA duplex stability in diverse molecular crowding conditions

GC Content Varies Enormously Across Life

Not all genomes use the four bases in equal proportions. The percentage of G and C bases (GC content) swings wildly across species. Among microorganisms the range is enormous, from below 25% to above 75%. Animals and plants are more constrained, but there are still differences: monocot plants (grasses, lilies, palms) have higher GC content than mammals, which in turn are higher than dicot plants (most broadleaf flowering plants).9PubMed Central. Variation, Evolution, and Correlation Analysis of C+G Content and Genome or Chromosome Size in Different Kingdoms and Phyla

Why GC content differs so much is an active research question. Thermal adaptation explains part of the pattern in heat-loving bacteria, but that cannot account for the variation in organisms that all live at moderate temperatures. One process that nudges genomes toward higher GC content is biased gene conversion, a side effect of DNA repair during recombination that tends to favor G and C bases over A and T. Modeling suggests this bias can evolve and stabilize through natural selection, pushing genomes gradually toward more GC-rich compositions.10Oxford Academic. Evolution of GC-biased gene conversion by natural selection

When the Rules Break: Hoogsteen Pairs

The A-T and G-C pairs described so far are Watson-Crick pairs, the standard geometry. But DNA bases can also form an alternative arrangement called Hoogsteen pairing, where one base flips around its bond to the sugar backbone and the two bases meet at different atoms. Hoogsteen pairs were once thought to be oddities limited to unusual DNA structures, but NMR studies have shown that base pairs in ordinary duplex DNA constantly flicker between Watson-Crick and Hoogsteen forms, existing as a dynamic equilibrium.11Nature. Transient Hoogsteen base pairs in canonical duplex DNA

These transient Hoogsteen pairs are short-lived and present at low abundance, roughly half a percent of all base pairs in solution at any given moment. But they are not randomly distributed. A structural survey of crystal structures found Hoogsteen pairs at about 0.3% of all base pairs, and A-T pairs were about four times more likely to adopt the Hoogsteen geometry than G-C pairs.12Nucleic Acids Research. New insights into Hoogsteen base pairs in DNA duplexes from a structure-based survey Hoogsteen pairs also cause measurable DNA bending of about 14 degrees per pair, which may play a role in how proteins recognize specific DNA sequences.

Hoogsteen pairing also shows up in less benign contexts. It appears frequently in DNA bound to certain antibiotics, in damaged DNA, and at the active sites of polymerases that replicate past damaged sites.13PubMed Central. A historical account of Hoogsteen base-pairs in duplex DNA Far from being a curiosity, the ability to switch between Watson-Crick and Hoogsteen geometries appears to be a built-in feature that expands what DNA can do.

Tautomers, Mispairing, and Spontaneous Mutations

Watson and Crick themselves pointed out a potential vulnerability in the pairing system. Each base normally exists in one dominant chemical form, but it can briefly shift to a rare alternative form called a tautomer, where a hydrogen atom moves to a different position on the ring. When this happens during DNA replication, the tautomeric base can form hydrogen bonds with the wrong partner, and the resulting mispair looks, geometrically, almost identical to a correct Watson-Crick pair.14PubMed Central. Structural evidence for the rare tautomer hypothesis of spontaneous mutagenesis

This is one of the main sources of spontaneous point mutations. The polymerase enzyme cannot tell that anything is wrong because the mispair has the right shape, so it incorporates the wrong base. Once the tautomer relaxes back to its normal form, the mismatch becomes obvious, but by then the error may already be locked in. Structural studies have confirmed that these tautomeric mismatches genuinely mimic canonical base pairs inside the polymerase active site.15PubMed Central. Structural Insights Into Tautomeric Dynamics in Nucleic Acids and in Antiviral Nucleoside Analogs Other mismatch mechanisms, including wobble pairing and Hoogsteen pairing, also contribute to replication errors, but the tautomer route is the one Watson and Crick originally predicted, and it took decades of structural work to prove they were right.16PubMed Central. The influence of base pair tautomerism on single point mutations in aqueous DNA

Oxidative Damage and Misreading

Beyond tautomeric shifts, chemical damage to bases can also disrupt normal pairing. The most common form of oxidative DNA damage produces a modified base called 8-oxoguanine, which is just guanine with an extra oxygen atom. This small change has a big consequence: 8-oxoguanine prefers to flip into a different orientation and pair with adenine instead of its usual partner cytosine. Left unrepaired, this leads to a specific mutation pattern where what was originally a G-C pair becomes a T-A pair in subsequent rounds of replication.17PubMed Central. Dynamic behavior of DNA base pairs containing 8-oxoguanine Cells have dedicated repair enzymes to find and fix these lesions, but some inevitably slip through, and 8-oxoguanine damage accumulates with age.

Why RNA Uses Uracil Instead of Thymine

If you have encountered RNA, you know it uses uracil (U) where DNA uses thymine (T). Uracil and thymine are nearly identical molecules; thymine is just uracil with an extra methyl group. Both pair with adenine using two hydrogen bonds, so the pairing rule is preserved. The reason DNA uses the methylated version likely comes down to damage management.

A recent study compared the photodamage pathways of the two bases under ultraviolet light similar to what reached Earth’s surface before the ozone layer formed. Thymine is actually more photoreactive overall than uracil, absorbing more UV. But it channels that damage preferentially toward a type of lesion (cyclobutane pyrimidine dimers) that is reversible and repairable, while uracil produces more of the irreversible type of damage.18PubMed Central. UV photodamage pathways and the evolutionary selection of thymine over uracil in early genetic systems The selection pressure, in other words, was not to avoid damage entirely but to steer damage toward pathways the cell could fix. DNA, as the permanent genomic record, benefited from this extra layer of protection. RNA, which is short-lived and disposable, gets by without it.

Methylation and Epigenetic Marks on CG Pairs

The C in a CG pair has a special property that none of the other bases share in the same way: it is the primary target for DNA methylation, one of the cell’s main epigenetic switches. Enzymes can attach a methyl group to the 5-carbon position of cytosine, converting it to 5-methylcytosine. This modification does not change the base pairing at all. Crystal structures of methylated DNA show that the overall helix shape, thermodynamic behavior, and base-pair geometry remain essentially the same as unmodified DNA, and polymerases cannot distinguish 5-methylcytosine from regular cytosine during replication.19Nucleic Acids Research. Crystal structures of B-DNA dodecamer containing the epigenetic modifications 5-hydroxymethylcytosine or 5-methylcytosine

What methylation does change is the physical behavior of the DNA at a finer scale. The methyl groups alter stacking interactions between adjacent CG steps, inhibiting overtwisting and softening certain mechanical modes in the helix.20PubMed Central. 5-Methylation of cytosine in CG:CG base-pair steps: a physicochemical mechanism for the epigenetic control of DNA nanomechanics Even more dramatically, methylation affects how easily the two strands re-form after being separated. Experiments using magnetic tweezers showed that methylated DNA was slightly harder to unzip but dramatically harder to rezip, with the hysteresis between unzipping and rezipping increasing over six-fold compared to unmethylated DNA.21Nucleic Acids Research. 5-Methyl-cytosine stabilizes DNA but hinders DNA hybridization revealed by magnetic tweezers and simulations These mechanical effects help explain how a tiny chemical tag on a CG pair can silence a gene without changing a single letter of the genetic code.

Expanding the Alphabet With Synthetic Base Pairs

Researchers have asked: can you go beyond A-T and G-C? The answer is yes. Several artificial base pairs have been designed that can sit alongside the natural ones inside a DNA helix, be copied by polymerases during PCR amplification, and even be transcribed into RNA.22PubMed Central. Unnatural base pair systems toward the expansion of the genetic alphabet in the central dogma Unlike the natural pairs, which rely on hydrogen bonds for specificity, some of the most successful synthetic pairs work through hydrophobic interactions, essentially two greasy surfaces that prefer each other’s company over the watery environment.

One pair discovered through large-scale screening, called d5SICS:dMMO2, was identified from a pool of 3,600 candidates as the combination best recognized by a natural DNA polymerase, and it was subsequently optimized for efficiency and selectivity.23PubMed Central. Discovery, characterization, and optimization of an unnatural base pair for expansion of the genetic alphabet Expanding the genetic alphabet this way could allow cells to produce proteins containing unnatural amino acids, opening doors for drug design and materials science. The work also underscores something about the natural system: A-T and G-C are not the only chemically possible pairs, but they are an exceptionally good solution to the problem of storing and copying information reliably at biological temperatures.

DNA Nanotechnology and the Predictability of Pairing

The strict specificity of A-T and G-C pairing has been exploited far beyond biology. In structural DNA nanotechnology, researchers design short DNA strands whose sequences are chosen so that the strands self-assemble into predetermined shapes: cubes, polyhedra, flat tiles, and intricate three-dimensional structures with features measured in nanometers. The entire field rests on the predictability of Watson-Crick pairing: if you know the sequences, you know which strands will bind where and in what orientation.24PubMed Central. Structural DNA Nanotechnology: State of the Art and Future Perspective DNA origami, one of the best-known techniques, uses a long scaffold strand and hundreds of short “staple” strands that fold the scaffold into a target shape purely through programmed base pairing. No glue, no enzymes, just A matching T and C matching G, over and over, until the desired architecture emerges in a test tube.