In nearly all living organisms, the codon AUG (spelled ATG in DNA) serves as the start signal for building a protein, while three codons act as stop signals: UAA, UAG, and UGA. These signals are triplets of nucleotide bases in messenger RNA that the ribosome reads like punctuation marks, telling it where to begin and where to end. The real story, though, is messier than the textbook version suggests, because nature bends and sometimes breaks these rules in ways that matter for medicine, evolution, and the frontiers of genetic engineering.
The Start Codon and What Makes It Special
AUG is the near-universal start codon. When a ribosome encounters AUG in the right context, it recruits a special initiator transfer RNA carrying the amino acid methionine, and protein synthesis begins. Every protein your cells produce starts with methionine at the outset (though it is often clipped off later). This is true across bacteria, animals, plants, and fungi.
The word “context” matters here. AUG appears many times throughout a typical messenger RNA molecule, but the ribosome does not start translating at just any AUG it bumps into. In human cells and other eukaryotes, the surrounding sequence called the Kozak sequence helps the ribosome recognize the correct AUG. The nucleotides immediately upstream of the start codon, especially at the minus-three position, play a critical role. A recent cryo-electron microscopy study showed exactly how this works at the atomic level: the three nucleotides just before AUG slot into a pre-formed clamp in the ribosome’s small subunit, and a specific amino acid on the initiation factor eIF2α forms a stacking interaction with a purine base at the minus-three position. A purine (A or G) at that spot stabilizes the complex; a smaller pyrimidine base (C or U) cannot form the same interaction, weakening start-site recognition and making translation less efficient.
1PubMed Central. Translation initiation by the Kozak mRNA sequence is based on a conformational readout on the ribosomeBacteria use a different system. Many bacterial messenger RNAs carry an upstream sequence called the Shine-Dalgarno sequence that pairs with the ribosome’s own RNA to position the start codon correctly. Research in E. coli has shown that while Shine-Dalgarno sequences are common, other guanine- and uracil-rich motifs also help ribosomes find the start site, likely by increasing the local concentration of ribosomes near the start codon rather than through precise base-pairing.
2PubMed Central. Evidence for context-dependent complementarity of non-Shine-Dalgarno ribosome binding sites to Escherichia coli rRNAThe Three Stop Codons
UAA, UAG, and UGA are the three stop codons, sometimes called “nonsense codons” because they do not code for any amino acid. Instead, they signal the ribosome to release the finished protein and disassemble. Each has a traditional nickname: UAA is “ochre,” UAG is “amber,” and UGA is “opal” (or sometimes “umber”). These names come from inside jokes among the molecular biologists who discovered them in the 1960s and have stuck around ever since.
Unlike start codons, stop codons are not recognized by transfer RNAs. Instead, proteins called release factors read the stop signal and trigger the release of the newly made protein. In bacteria, two release factors split the work: RF1 recognizes UAA and UAG, while RF2 recognizes UAA and UGA. A third factor, RF3, helps the others detach from the ribosome after they’ve done their job.3PubMed Central. Distinct roles for release factor 1 and release factor 2 in translational quality control In human cells and other eukaryotes, a single release factor called eRF1 recognizes all three stop codons. Structural studies have shown that eRF1 physically mimics the shape of a transfer RNA molecule, allowing it to fit into the same slot on the ribosome where a tRNA would normally go.4Cell. The Crystal Structure of Human Eukaryotic Release Factor eRF1—Mechanism of Stop Codon Recognition and Peptidyl-tRNA Hydrolysis A partner factor, eRF3, uses energy from GTP to coordinate the process and ensure the protein is properly released.5PubMed Central. Cryo-EM structure of the mammalian eukaryotic release factor eRF1-eRF3-associated termination complex
High-resolution imaging has revealed exactly how eRF1 reads the stop codon at the molecular level. When a stop codon sits in the ribosome’s reading site, a nucleotide in the ribosomal RNA flips into position and stacks against the second and third bases of the stop codon. This rearrangement also pulls the nucleotide immediately after the stop codon into the reading site, where it gets pinned in place by another ribosomal RNA residue. The identity of that fourth-position nucleotide can influence how efficiently termination occurs, which is one reason stop codon context matters and not all stop codons work equally well.6PubMed Central. Structural basis for stop codon recognition in eukaryotes
Non-Canonical Start Codons
AUG gets all the attention, but cells can start translation from other codons too. Near-cognate codons like CUG, GUG, and UUG, which differ from AUG by a single base, function as start codons in certain contexts, especially in bacteria. In E. coli, researchers systematically tested all 64 possible codons for their ability to initiate translation. They found measurable initiation from many non-canonical codons, with single-base mismatches from AUG initiating at roughly 0.01 to 3% the rate of AUG, and codons with multiple mismatches at even lower rates.7Nucleic Acids Research. Measurements of translation initiation from all 64 codons in E. coli
These numbers might sound trivially small, but they have real biological consequences. Ribosome profiling and proteomic studies have uncovered thousands of previously unknown short coding sequences in both bacterial and eukaryotic genomes, many of which begin at non-AUG codons. These small open reading frames were invisible to traditional gene-finding software precisely because the software assumed genes start with AUG.8PubMed Central. Non-AUG start codons: Expanding and regulating the small and alternative ORFeome Some of these non-AUG-initiated genes participate in important biological processes like stress responses and development, so the textbook rule that “AUG is the start codon” is a useful simplification rather than an absolute law.
When Stop Codons Code for Amino Acids
Two of the three stop codons have been co-opted in nature to encode unusual amino acids, expanding the genetic code beyond the standard twenty.
UGA, normally a stop signal, can instead direct the insertion of selenocysteine, sometimes called the “21st amino acid.” Selenocysteine contains the trace element selenium and is essential to a family of enzymes called selenoproteins, which play roles in antioxidant defense and thyroid hormone metabolism. For UGA to be read as selenocysteine rather than “stop,” the messenger RNA must contain a specific stem-loop structure called a SECIS element. In eukaryotes, this element sits in the untranslated region downstream of the coding sequence and recruits specialized factors that override the normal termination machinery.9PubMed Central. Novel structural determinants in human SECIS elements modulate the translational recoding of UGA as selenocysteine
UAG, the amber stop codon, serves a similar dual role in certain archaea and bacteria, where it encodes pyrrolysine, the “22nd amino acid.” Pyrrolysine is found in enzymes that break down methylamines, an environmentally significant metabolism in methane-producing archaea. In the model organism Methanosarcina acetivorans, UAG has genuine dual meaning: it encodes pyrrolysine at internal positions in genes that need it and functions as a stop signal elsewhere.10PubMed Central. Methanogenic archaea encoding Pyrrolysine maintain ambiguous amber codon usage Proteomic analysis of certain other archaea has confirmed that some species go even further, consistently incorporating pyrrolysine at all their TAG codons, effectively running an alternative genetic code.11PubMed. An archaeal genetic code with all TAG codons as pyrrolysine
Organisms That Reassigned All Three Stop Codons
The most extreme departure from the standard genetic code has been found in certain single-celled organisms called ciliates. The ciliate Condylostoma magnum uses all three traditional stop codons to encode amino acids: UAA and UAG code for glutamine, and UGA codes for tryptophan. Since every one of the 64 possible codons now specifies an amino acid, the question becomes: how does this organism know when to stop making a protein?12Cell. Genetic Codes with No Dedicated Stop Codon: Context-Dependent Translation Termination
The answer is position. Evidence suggests that in Condylostoma magnum, these codons are decoded as amino acids when they appear in the interior of a gene but trigger translation termination when they occur near the physical end of the messenger RNA.13PubMed Central. Novel Ciliate Genetic Code Variants Including the Reassignment of All Three Stop Codons to Sense Codons in Condylostoma magnum This context-dependent system is a striking example of how the genetic code, often presented as universal and frozen, has continued to evolve in certain lineages. A related unclassified karyorelict ciliate uses the same scheme, so this is not a one-off oddity.
Natural Readthrough of Stop Codons
Even in organisms with a perfectly standard genetic code, stop codons are not always obeyed. In a process called translational readthrough, the ribosome occasionally ignores a stop codon and inserts an amino acid instead, continuing to translate the messenger RNA beyond the normal endpoint. When this happens by chance it is usually harmless, but in some genes it has been selected by evolution as a deliberate regulatory trick. This functional translational readthrough creates extended versions of proteins with additional domains, effectively getting two protein variants from a single gene.14PubMed Central. Functional Translational Readthrough: A Systems Biology Perspective
The rate of readthrough can vary dramatically between tissues. In fruit flies, researchers studying a transcription factor gene called traffic jam found that readthrough levels ranged from minimal in some tissues to about 20% in the head. The variation correlated with the abundance of a specific splice variant of the release factor eRF1, and this particular variant was most common in nervous system tissues. Overexpression of the variant in cultured cells boosted readthrough, particularly at UGA stop codons.15Nucleic Acids Research. Tissue-specific regulation of translational readthrough tunes functions of the traffic jam transcription factor This suggests that cells can fine-tune how strictly they obey stop codons depending on which tissue they belong to, adding a layer of gene regulation that doesn’t involve changing the DNA sequence at all.
When Stop Codons Appear Too Early
A mutation that converts an amino acid codon into a premature stop codon, called a nonsense mutation, is one of the most damaging things that can happen to a gene. The resulting truncated protein is usually non-functional. Cells have a surveillance system called nonsense-mediated mRNA decay that detects messenger RNAs containing premature stop codons and destroys them before they can produce much defective protein.16PubMed Central. Nonsense-Mediated mRNA Decay, a Finely Regulated Mechanism This quality control pathway prevents the cell from accumulating harmful protein fragments, but it also means the gene produces little or no protein at all, which can itself cause disease.
When ribosomes do stall or encounter aberrant termination events, eukaryotic cells have an additional cleanup system. Stalled ribosomes can split, leaving the large subunit still attached to the incomplete protein chain. A dedicated quality-control pathway then tags the faulty protein with ubiquitin, marking it for destruction by the cell’s protein-recycling machinery.17PubMed Central. Distinct types of translation termination generate substrates for ribosome-associated quality control
Drugs That Force the Ribosome to Read Through Premature Stop Codons
Premature stop codons are estimated to cause roughly 10 to 20% of inherited genetic diseases, including forms of cystic fibrosis, Duchenne muscular dystrophy, and certain cancers where tumor suppressor genes are knocked out.18PubMed Central. Genome-scale quantification and prediction of pathogenic stop codon readthrough by small molecules One therapeutic strategy is to use drugs that make the ribosome skip over the premature stop codon and produce a full-length (or near-full-length) protein instead. Over 50 small molecules have been identified with this readthrough-promoting ability, falling into two broad categories: aminoglycoside antibiotics and non-aminoglycoside compounds.19PubMed Central. Pharmaceuticals Promoting Premature Termination Codon Readthrough: Progress in Development
Aminoglycosides like gentamicin were the first drugs found to promote readthrough, but they cause kidney and hearing damage at the doses needed, so the search for safer alternatives has been intense. Combination approaches look promising. One study found a compound called CDX5-1 that had no readthrough activity on its own but boosted the effect of the aminoglycoside G418 by up to 180-fold when the two were used together. The combination worked across multiple nonsense mutations in different genes, including the tumor suppressor TP53 and genes responsible for rare diseases like neuronal ceroid lipofuscinosis and Duchenne muscular dystrophy.20Nucleic Acids Research. Novel small molecules potentiate premature termination codon readthrough by aminoglycosides
A major obstacle is that not all premature stop codons respond equally well to drugs. The identity of the stop codon itself, the surrounding sequence context, and the position within the gene all influence how effectively a drug can force readthrough. A genome-scale study quantified readthrough of roughly 5,800 human pathogenic stop codons by eight different drugs, providing the first systematic map of which mutations are most druggable and which resist current compounds.18PubMed Central. Genome-scale quantification and prediction of pathogenic stop codon readthrough by small molecules This kind of data is essential for figuring out which patients might benefit from readthrough therapy and which will not.
Rewriting the Genetic Code in the Lab
Synthetic biologists have taken the natural flexibility of start and stop codons to an extreme by engineering organisms with radically altered genetic codes. In one landmark project, researchers replaced all 1,195 TGA stop codons in an E. coli genome with TAA, then engineered the bacterium’s release factor and a transfer RNA so that UGA and UAG were no longer recognized as stop signals at all. The result was an organism that uses UAA as its sole stop codon, freeing up UAG and UGA to encode two different non-standard amino acids in the same protein with over 99% accuracy.21PubMed Central. Engineering a genomically recoded organism with one stop codon
This kind of genome recoding goes beyond swapping stop codons. Another project built a synthetic E. coli genome that uses only 57 of the 64 possible codons to encode all its proteins, eliminating seven codons entirely. The effort required overcoming the lethal effects of over 62,000 synonymous codon swaps and more than 11,000 additional genomic edits.22PubMed Central. Synthetic genomes unveil the effects of synonymous recoding Such recoded organisms have practical applications: because their genetic code is incompatible with natural organisms, they are resistant to viral infection (viruses cannot translate their own genes properly inside a recoded host), and they cannot exchange genetic material with wild bacteria, providing a built-in biocontainment mechanism.
How Gene-Finding Software Uses Stop Codons
Start and stop codons play a practical role well beyond the cell. In bioinformatics, identifying potential genes in a newly sequenced genome depends heavily on finding open reading frames, stretches of DNA that begin with a start codon and end with a stop codon without any intervening stop codons in the same reading frame. Gene-finding algorithms use the frequency and distribution of start and stop codons alongside other signals like codon usage patterns and promoter motifs to predict where genes are located.23Gene. GC content dependency of open reading frame prediction via stop codon frequencies
The expected frequency of stop codons in random DNA depends on the organism’s GC content. Organisms with DNA that is rich in G and C bases will naturally have fewer occurrences of the AT-rich stop codons TAA and TAG, which means longer stretches of DNA will appear gene-like simply by chance. This makes gene prediction more challenging in GC-rich genomes and is one reason why computational gene prediction for organisms with unusual base compositions sometimes produces unreliable results. The discovery of organisms like Condylostoma magnum, where stop codons mean different things in different positions, adds another wrinkle: standard gene-finding tools that treat stop codons as absolute endpoints would misannotate genes in these organisms entirely.