What Are the Base Pairs for RNA and How Do They Function?

RNA uses four main bases that pair in two canonical combinations: adenine (A) pairs with uracil (U), and guanine (G) pairs with cytosine (C). These pairings follow the same hydrogen-bonding logic as DNA’s famous ladder, with one critical swap: RNA carries uracil where DNA carries thymine. But unlike DNA, which primarily exists as a stable double helix, RNA folds back on itself into wildly diverse shapes, and that folding depends on a much richer pairing vocabulary than just A-U and G-C. Non-canonical pairs, chemical modifications, and long-range interactions all expand what base pairing means in the RNA world.

The Two Standard Pairs and Why RNA Uses Uracil

In a canonical RNA duplex, adenine forms two hydrogen bonds with uracil, and guanine forms three hydrogen bonds with cytosine. The G-C pair is therefore stronger than A-U, which has real consequences for the stability of any RNA structure: regions rich in G-C pairs tend to hold together more firmly, while A-U-heavy stretches are more easily unwound or opened up by the cell’s machinery.

The substitution of uracil for thymine is one of the oldest biochemical distinctions between RNA and DNA. Thymine is essentially uracil with an extra methyl group attached. That methyl group creates a stabilizing interaction when thymine stacks against adenine in a DNA double helix, helping to hold the structure rigid. In RNA, the absence of that methyl group in uracil contributes to the molecule’s flexibility and its tendency to form complex three-dimensional shapes rather than simple linear duplexes.1PubMed. An analysis of the different behavior of DNA and RNA through the study of the mutual relationship between stacking and hydrogen bonding Recent work has also shown that the choice between thymine and uracil may have been shaped early in evolution by how each base handles ultraviolet light damage. Thymine channels UV damage toward a type of lesion that can be reversed without enzymes, while uracil is more prone to forming irreversible damage products.2PubMed Central. UV photodamage pathways and the evolutionary selection of thymine over uracil in early genetic systems This gave DNA a stability advantage for long-term information storage, while RNA’s use of uracil was not a liability for a molecule whose lifespan in the cell is typically short.

The G-U Wobble Pair and Other Non-Canonical Pairings

If RNA relied solely on A-U and G-C, it could form straightforward helices and not much else. What makes RNA structurally versatile is that it tolerates, and even depends on, dozens of non-standard base pairings. The most important of these is the G-U wobble pair. G and U can form two hydrogen bonds in a slightly shifted geometry compared to a standard Watson-Crick pair, and this wobble pair turns up in nearly every major class of RNA across all domains of life.3PubMed Central. The G x U wobble base pair. A fundamental building block of RNA structure crucial to RNA function in diverse biological systems. Its thermodynamic stability is close to that of a Watson-Crick pair, so it can often substitute for G-C or A-U without destabilizing a structure. Yet it also has unique chemical properties that Watson-Crick pairs cannot replicate, making it functionally irreplaceable in certain contexts like tRNA recognition and RNA splicing.

Modeling studies of tandem G-U pairs have resolved a long-standing debate about their hydrogen bonding. When two G-U wobble pairs stack directly on top of each other in one orientation, the inner pair can be held together by just a single hydrogen bond, which is weaker than expected. But when G-U pairs sit at the ends of helices, they actually increase overall duplex stability, largely because their hydrogen bonding is stronger in that terminal position.4PubMed. Evaluating Hydrogen Bonds and Base Stacking of Single, Tandem and Terminal GU Mismatches in RNA with a Mesoscopic Model

Beyond the wobble pair, RNA bases can interact using three distinct edges: the Watson-Crick edge, the Hoogsteen edge (the side of a purine ring opposite the Watson-Crick face), and the Sugar edge (which includes the 2′-hydroxyl group unique to RNA’s ribose sugar). Because two bases can meet in either a cis or trans orientation relative to their backbone connections, this gives rise to twelve basic geometric families of base pairs.5PubMed Central. Geometric nomenclature and classification of RNA base pairs Computational analysis of high-resolution RNA crystal structures has catalogued over 150 distinct base pair types across these families, including pairs involving chemically modified bases.6PubMed Central. Estimating Strengths of Individual Hydrogen Bonds in RNA Base Pairs: Toward a Consensus between Different Computational Approaches This enormous pairing repertoire is what allows RNA to build the intricate folds needed for everything from catalysis to gene regulation.

How Base Pairs Build RNA Structure

RNA folds in a hierarchical way. First, stretches of complementary sequence pair up to form helices, loops, and bulges. These local elements make up what is called secondary structure, and they are much more stable thermodynamically than the longer-range contacts that come later. Because of this stability gap, secondary structure forms first and can be predicted with reasonable accuracy from sequence alone.7PubMed. How RNA folds Only after the secondary scaffold is in place do tertiary interactions kick in, folding the molecule into its final three-dimensional shape without drastically distorting the existing helices.

Bulges, where one or more unpaired nucleotides interrupt a helix, play a surprisingly active role. They generally destabilize a duplex, but the degree of destabilization depends on the stability of the surrounding helical stems.8PubMed Central. Influence of two bulge loops on the stability of RNA duplexes This is not just a structural curiosity. In the spliceosome, the molecular machine that edits messenger RNA, a dynamic bulge in a small RNA called U6 helps coordinate a magnesium ion that is essential for the first chemical step of RNA splicing.9PubMed Central. A dynamic bulge in the U6 RNA internal stem-loop functions in spliceosome assembly and activation A seemingly simple irregularity in base pairing turns out to be the hinge on which a major cellular process swings.

Pseudoknots and Tertiary Contacts

One of the most important tertiary structures RNA can form is the pseudoknot. A pseudoknot occurs when a loop in a stem-loop structure pairs with a sequence outside the loop, creating an interlocked arrangement of two helical segments.10PubMed Central. Structure and function of pseudoknots involved in gene expression control The result is compact and rigid. NMR studies have shown that in a classical pseudoknot, the loops crossing the grooves of the helices make specific hydrogen bonds, including contacts with highly conserved nucleotides, and the entire structure retains a degree of internal flexibility that may help it interact with proteins and other RNA molecules.11PubMed. NMR structure of a classical pseudoknot: interplay of single- and double-stranded RNA

Pseudoknots are not just architectural motifs. In many viruses, including SARS-CoV-2, a pseudoknot jams into the entrance of the ribosome’s mRNA channel during translation, creating mechanical tension that causes the ribosome to slip backward by one nucleotide on the mRNA. This “programmed ribosomal frameshifting” is essential for the virus to produce its RNA-copying enzyme.12PubMed Central. Structural basis of ribosomal frameshifting during translation of the SARS-CoV-2 RNA genome The tertiary base pairs within these pseudoknots, particularly base-triple interactions where loop nucleotides dock into the minor groove of an adjacent helix, directly influence how efficiently this frameshifting occurs. Replace the loop’s adenosines with pyrimidines, and frameshifting efficiency drops to roughly a tenth of its original level.13Nucleic Acids Research. Coordination among tertiary base pairs results in an efficient frameshift-stimulating RNA pseudoknot

Wobble Decoding in Translation

Every cell reads its messenger RNA three nucleotides at a time, each triplet (called a codon) specifying a particular amino acid. There are 61 codons that encode amino acids, yet cells get by with far fewer transfer RNA (tRNA) molecules to decode them. Francis Crick explained this mismatch in 1966 with the wobble hypothesis: the first position of the tRNA anticodon (position 34) does not have to follow strict Watson-Crick rules. A uridine at that position can pair not only with adenosine but also with guanosine, and inosine (a modified base) can pair with uridine, cytidine, and adenosine.14PubMed Central. Celebrating wobble decoding: Half a century and still much is new This expanded pairing at the wobble position is why the genetic code is degenerate: multiple codons can specify the same amino acid because a single tRNA can read more than one of them.

Wobble decoding is not an error-prone shortcut. The ribosome actively checks the geometry of the codon-anticodon pair, and wobble pairs like G-U fit close enough to Watson-Crick dimensions that they pass inspection. This balance between flexibility and accuracy is one of the reasons translation can be simultaneously fast and reliable.

Catalytic RNA and the Role of Base Pairing in Ribozymes

Some RNA molecules catalyze chemical reactions, functioning as enzymes without any protein component. These ribozymes depend on precise base-pairing networks to position their active sites. The hammerhead ribozyme, one of the smallest known, uses tertiary interactions far from its cleavage site to arrange two critical guanine residues in positions consistent with acid-base catalysis.15PubMed Central. Tertiary contacts distant from the active site prime a ribozyme for catalysis Without those distant base-pairing contacts, the active site does not adopt the right shape and catalysis stalls.

Systematic mutagenesis experiments across several self-cleaving ribozymes have confirmed a clear trend: mutations that disrupt two base pairs at once produce the most damaging combined effect, while double mutations that happen to create a new Watson-Crick or G-U wobble pair can actually rescue activity.16eLife. RNA sequence to structure analysis from comprehensive pairwise mutagenesis of multiple self-cleaving ribozymes The message is that base pairing is not just a scaffold that holds the ribozyme together passively. It is an active participant in catalysis, tuning geometry at angstrom-level precision.

The Tetrahymena group I intron ribozyme provides an especially striking example. Its crystal structure reveals a “triple-helical sandwich” at the active site where the attacking guanosine forms a coplanar base triple with a G-C pair, and this triple is flanked above and below by additional layers of base triples involving Hoogsteen and minor-groove contacts.17Molecular Cell. Crystal Structure of an Active Tetrahymena Ribozyme at 3.8 Ã… Resolution These stacked triples lock the substrate in place and explain why single-nucleotide mutations at these positions are lethal to the enzyme’s function.

Chemical Modifications That Alter Base Pairing

Cells do not leave their RNA bases untouched after transcription. Over 170 distinct chemical modifications have been identified in RNA, and several directly change how bases pair with each other. Two of the most consequential are adenosine-to-inosine editing and uridine-to-pseudouridine conversion.

Adenosine-to-inosine (A-to-I) editing is carried out by ADAR enzymes, which chemically convert adenosine into inosine. The key consequence is that inosine pairs preferentially with cytidine rather than uridine, effectively changing what the ribosome reads at that position. When this happens inside a protein-coding region, it can swap one amino acid for another, diversifying the proteins a cell produces from a single gene. A large-scale analysis of over 9,000 human tissue samples identified more than 1,500 such recoding sites in the human transcriptome.18Nature Communications. Landscape of adenosine-to-inosine RNA recoding across human tissues Beyond recoding, inosine modification also influences RNA splicing, stability, and protein binding.19RNA. Structural and functional effects of inosine modification in mRNA

Pseudouridine is the most abundant internal modification in RNA. It is an isomer of uridine in which the base is attached to the sugar through a carbon-carbon bond instead of the usual carbon-nitrogen bond. This seemingly small chemical change gives pseudouridine an extra hydrogen-bond donor, enabling it to form stable pairs with all four canonical bases, not just adenine. Thermodynamic measurements show that replacing uridine with pseudouridine can stabilize an RNA duplex, though the degree of stabilization depends on where in the helix the substitution occurs and what its neighbors are.20PubMed Central. The contribution of pseudouridine to stabilities and structure of RNAs Detailed molecular dynamics work confirms this position-dependence: the same substitution can be destabilizing in one context and globally stabilizing in another.21PubMed Central. Structural and dynamic effects of pseudouridine modifications on noncanonical interactions in RNA This sensitivity to context is a reminder that base pairing in RNA is not just about the identity of two facing nucleotides; the surrounding sequence and structure matter enormously.

G-Quadruplexes and Guanine’s Unusual Arrangements

Not all RNA base pairing is edge-to-edge between two partners. In guanine-rich sequences, four guanines can arrange themselves into a flat square called a G-quartet, held together by a distinctive pattern of hydrogen bonds where each guanine simultaneously donates and accepts hydrogen bonds with its neighbors. When two or more of these quartets stack on top of each other, they form a G-quadruplex (G4), an extremely stable structure.22PubMed Central. RNA G-Quadruplexes in Biology: Principles and Molecular Mechanisms RNA G-quadruplexes tend to be more stable than their DNA counterparts, largely because the extra 2′-hydroxyl group on RNA’s ribose sugar provides additional favorable electrostatic interactions.23PubMed Central. RNA versus DNA G-Quadruplex: The Origin of Increased Stability

Finding these structures in biologically important sequences has been more challenging than anticipated. Recent experimental work showed that most candidate sequences either failed to form a quadruplex at all or formed intermolecular (multi-strand) versions instead of the intramolecular fold that would be relevant inside a cell. Unambiguous two-tetrad intramolecular G4 structures were rare, and experimental detection was complicated by the fact that both standard assays and fluorescent probes can give misleading signals for RNA.24Nucleic Acids Research. RNA G-quadruplex formation in biologically important transcribed regions: can two-tetrad intramolecular RNA quadruplexes be formed? The biology of RNA G-quadruplexes is real, but claims about how often they form in living cells should be taken with a grain of salt.

Long-Distance Base Pairing in Viral RNA

In most textbook depictions, base pairing happens between nucleotides that are close together in sequence. But some viral RNAs exploit base pairing between partners separated by thousands of nucleotides. In Barley yellow dwarf virus, a loop structure near the ribosomal frameshifting site must pair with a complementary sequence located about four kilobases downstream. When researchers introduced mismatches into either the loop or the distant bulge, frameshifting was abolished and viral replication in plant cells shut down entirely. Restoring complementarity between the two sites rescued both functions.25PubMed Central. A -1 ribosomal frameshift element that requires base pairing across four kilobases suggests a mechanism of regulating ribosome and replicase traffic on a viral RNA The model that emerged is that the virus uses these long-distance base-pairing events to coordinate translation and replication on the same RNA molecule, essentially switching the RNA from a ribosome-occupied state to a replication-ready state.

This kind of architectural trick is not unique to plant viruses. Coronaviruses, retroviruses, and many other RNA viruses rely on carefully tuned secondary and tertiary structures within their genomes to regulate when and how their genes are read. The base pairs holding these structures together are potential drug targets: disrupt the right pair and the virus cannot properly express its genes.

Therapeutic Applications Built on Base Pairing

Modern RNA-targeting drugs are designed around Watson-Crick pairing rules. Antisense oligonucleotides (ASOs) are short synthetic strands engineered to pair with a specific stretch of target RNA. Once bound, they can trigger destruction of the target by cellular enzymes, block translation, or alter splicing patterns.26PubMed. RNA targeting therapeutics: molecular mechanisms of antisense oligonucleotides as a therapeutic platform Small interfering RNAs (siRNAs) work on a similar principle, using a short guide strand that pairs with the target mRNA and recruits a protein complex to slice it apart.

Specificity is the central design challenge. The factors that make an ASO bind tightly to the intended target are exactly the same factors that can cause it to bind unintended targets elsewhere in the transcriptome.27Nucleic Acids Research. Managing the sequence-specificity of antisense oligonucleotides in drug discovery Off-target pairing has the same chemistry as on-target pairing; the only defense is careful sequence selection, chemical modifications to the backbone that raise the energetic bar for mismatched binding, and thorough screening.

Probing RNA Structure in Living Cells

Understanding which bases are paired and which are free is essential for deciphering RNA function, but RNA folds differently inside a cell than in a test tube. A technique called SHAPE-MaP (selective 2′-hydroxyl acylation analyzed by primer extension and mutational profiling) lets researchers measure the flexibility of every nucleotide in an RNA molecule at single-nucleotide resolution, even inside living cells. Small chemical probes react with the 2′-hydroxyl group of flexible, unpaired nucleotides but leave paired nucleotides untouched. The modified positions are then read out through sequencing.28PubMed Central. In-cell RNA structure probing with SHAPE-MaP A complementary approach uses dimethyl sulfate (DMS), which modifies unpaired adenines and cytosines, providing a different structural fingerprint that can be cross-checked against SHAPE results.29PubMed. Chemical Probing of RNA Structure In Vivo Using SHAPE-MaP and DMS-MaP Together, these methods have revealed that many RNAs adopt different structures in the cell than those predicted computationally, underscoring that base pairing is influenced by the crowded, protein-rich cellular environment.

Expanding the Alphabet

Researchers have pushed beyond nature’s four RNA bases by creating synthetic base pairs that can be replicated by DNA polymerase, transcribed into RNA, and potentially decoded by the ribosome. One well-characterized pair, called d5SICS-dNaM, can be copied and transcribed with fidelities sufficient for use as part of an expanded genetic alphabet.30PubMed Central. Transcription of an expanded genetic alphabet Other unnatural pairs have been engineered to function alongside the standard A-T and G-C pairs during PCR amplification, with the resulting synthetic DNA then successfully transcribed into RNA containing the new bases.31PubMed Central. Unnatural base pair systems toward the expansion of the genetic alphabet in the central dogma

Separately, experiments modeling prebiotic chemistry have shown that template-directed RNA copying using all four natural bases can proceed without enzymes when helped along by short activated oligonucleotide helpers, reaching fidelities around 98% and producing full-length products in good yield.32PubMed Central. Nonenzymatic copying of RNA templates containing all four letters is catalyzed by activated oligonucleotides These results strengthen the case that base pairing was robust enough to sustain information transfer in the RNA world, long before protein enzymes or DNA genomes existed. The same hydrogen-bonding logic that lets your cells read messenger RNA today was, in all likelihood, the very first molecular trick life learned.