Selenocysteine is an amino acid built around the trace element selenium, and it earns its “21st amino acid” title because cells make it and insert it into proteins using the genetic code itself, not by modifying a protein after it is already built. The standard genetic code accounts for 20 amino acids. Selenocysteine broke that count by co-opting a stop codon, UGA, and repurposing it as a signal to install a selenium-containing residue during protein synthesis. That repurposing requires an elaborate set of molecular machinery found across all three domains of life, from bacteria to humans, making selenocysteine far more than a biochemical curiosity.
How Selenocysteine Was Identified
The story begins with selenium itself, which was long known to be essential in trace amounts for animal health without anyone understanding why. The breakthrough came in 1986, when two research groups independently noticed something strange: the genes for two well-known selenium-containing enzymes, glutathione peroxidase and bacterial formate dehydrogenase, each had a UGA codon sitting right in the middle of their coding sequence. UGA normally tells the ribosome to stop translating, yet these genes clearly produced full-length proteins. It turned out that the UGA was not a stop signal at all in these contexts. Instead, it directed the cell to insert selenocysteine.1PubMed. Selenocysteine: the 21st amino acid
The discovery meant the genetic code was not as closed as textbooks had long claimed. A transfer RNA that had initially been mistaken for a quirky serine-carrying molecule turned out to be the dedicated carrier for selenocysteine. It was the only known amino acid in eukaryotes that is actually assembled while still attached to its own tRNA, rather than being made as a free molecule and then loaded on.2Advances in Nutrition. Biosynthesis of Selenocysteine, the 21st Amino Acid in the Genetic Code, and a Novel Pathway for Cysteine Biosynthesis
What Makes It Chemically Special
If you picture cysteine, one of the standard 20 amino acids, and swap its sulfur atom for selenium, you get selenocysteine. That single-element substitution has outsized chemical consequences. Selenium sits one row below sulfur on the periodic table, giving selenocysteine a larger, more polarizable reactive group. Its selenium-hydrogen bond breaks much more easily than cysteine’s sulfur-hydrogen bond, which means the selenocysteine side chain is already in its reactive, negatively charged form at the near-neutral pH found inside cells. Cysteine, by contrast, needs more alkaline conditions to reach the same state.3PubMed. Comparison of the Chemical Properties of Selenocysteine and Selenocystine with Their Sulfur Analogs
The practical upshot is speed. In simple chemical reactions at physiological pH, selenium-containing molecules can react up to ten thousand times faster than their sulfur counterparts, depending on the type of reaction.4PubMed. Why do proteins use selenocysteine instead of cysteine? Selenocysteine also has a much lower reduction potential, meaning it gives up electrons more readily. That property makes it ideal for enzymes that shuttle electrons around, since the selenium atom can cycle between oxidized and reduced states without getting permanently damaged. When cysteine is oxidized by a single electron, the resulting radical can itself attack and damage other parts of the protein; the equivalent selenium radical is far less destructive.4PubMed. Why do proteins use selenocysteine instead of cysteine?
The Molecular Machinery That Installs It
Getting selenocysteine into a protein is not simple. The cell has to distinguish between a UGA that means “stop translating” and a UGA that means “insert selenocysteine here.” In bacteria, the distinction comes from a structured loop of RNA called a SECIS element, located in the messenger RNA just downstream of the UGA codon. In eukaryotes and archaea, the SECIS element sits in the untranslated region at the far end of the messenger RNA, and the structural details differ considerably between domains of life even though the function is the same.5PubMed Central. Asgard archaeal selenoproteome reveals a roadmap for the archaea-to-eukaryote transition of selenocysteine incorporation machinery
Beyond the SECIS element, the cell needs a specialized elongation factor. Normal amino acids are ferried to the ribosome by a general-purpose delivery protein, but selenocysteine gets its own dedicated escort. In bacteria it is called SelB; in eukaryotes the equivalent is eEFSec. This factor forms a complex with the charged selenocysteine-tRNA and delivers it specifically to UGA codons that have the SECIS context. A separate protein, SBP2 in eukaryotes, binds the SECIS element and helps coordinate the whole process.6PubMed Central. Threading the needle: getting selenocysteine into proteins
The competition between stop-signal recognition and selenocysteine insertion is real. In a reconstituted bacterial system, about 35 to 40 percent of ribosomes successfully incorporated selenocysteine at a given UGA site. The release factor that normally triggers chain termination at UGA codons could still bind, but it only managed to end translation on ribosomes where selenocysteine insertion had already failed. In other words, the selenocysteine machinery wins the race when it is present, and the release factor sweeps up the leftovers.7PubMed Central. Partitioning between recoding and termination at a stop codon–selenocysteine insertion sequence
How Selenocysteine Is Built on Its tRNA
Unlike any other amino acid used in protein synthesis, selenocysteine does not exist as a free molecule waiting to be loaded onto a tRNA. Instead, the cell builds it in stages directly on the tRNA itself. The process starts when a standard enzyme loads serine onto the selenocysteine-specific tRNA. Then a kinase adds a phosphate group to the serine, creating phosphoseryl-tRNA. Finally, a specialized synthase swaps the phosphate-serine for selenocysteine, using a selenium donor called selenophosphate to provide the selenium atom.8PubMed Central. The canonical pathway for selenocysteine insertion is dispensable in Trypanosomes The active selenium donor in eukaryotes was confirmed to be selenophosphate through experiments showing that the final enzyme could produce selenocysteine only when selenophosphate was added to the reaction.9PLoS Biology. Biosynthesis of Selenocysteine on Its tRNA in Eukaryotes
This on-tRNA assembly line is one reason selenocysteine insertion is so tightly regulated. The cell can control how much selenocysteine gets made by adjusting the supply of selenium and the availability of the tRNA precursors. When selenium is scarce, less charged selenocysteine-tRNA is produced, and more UGA codons are read as stop signals rather than selenocysteine codons. That built-in sensitivity links selenoprotein production directly to the body’s selenium status.10PubMed Central. A quantitative model for the rate-limiting process of UGA alternative assignments to stop and selenocysteine codons
What Selenoproteins Actually Do in Your Body
Humans have about 25 genes encoding selenoproteins, and most of them are enzymes that deal with oxidative stress or redox chemistry. The best-known family is the glutathione peroxidases, which protect cells by neutralizing hydrogen peroxide and lipid hydroperoxides. The selenocysteine in their active site is what makes them catalytically effective; the selenium atom cycles between oxidized and reduced forms as the enzyme processes each molecule of peroxide.11Environmental Toxicology and Pharmacology. The biochemistry of selenium and the glutathione system
Another major group is the thioredoxin reductases, which maintain the cell’s internal redox balance. Mammalian thioredoxin reductase has a selenocysteine residue at its active site, and that residue is critical. When researchers replaced it with cysteine, the mutant enzyme retained only about 6 to 11 percent of normal catalytic activity, depending on the substrate tested.12PubMed. Mammalian thioredoxin reductase: oxidation of the C-terminal cysteine/selenocysteine active site forms a thioselenide, and replacement of selenium with sulfur markedly reduces catalytic activity These enzymes also play a role in redox signaling: when reactive oxygen species are generated inside a cell, they can oxidize the selenol group on thioredoxin reductase, temporarily reducing its activity and allowing downstream signaling molecules, including certain transcription factors and phosphatases, to change their behavior.13Journal of Biological Chemistry. Redox Regulation of Cell Signaling by Selenocysteine in Mammalian Thioredoxin Reductases
A third family with direct clinical relevance is the iodothyronine deiodinases, which control thyroid hormone activation. The thyroid gland mainly secretes thyroxine (T4), which is relatively inactive. Deiodinases convert T4 into T3, the form that actually drives metabolism in tissues. All three known deiodinase enzymes contain selenocysteine, and the selenium is essential for their activity.14PubMed. Local activation and inactivation of thyroid hormones: the deiodinase family The original cloning of the type I deiodinase revealed the UGA codon for selenocysteine in its gene, and this explained a longstanding puzzle: why experimental selenium deficiency impaired the body’s ability to convert T4 to T3.15PubMed. Type I iodothyronine deiodinase is a selenocysteine-containing enzyme
What Happens When the Machinery Breaks
Rare mutations in the genes that encode selenocysteine insertion machinery cause systemic selenoprotein deficiency, and the resulting clinical picture is remarkably broad. Mutations in the gene for SBP2 (the protein that binds the SECIS element in eukaryotes) impair the production of most or all selenoproteins and lead to a constellation of problems: abnormal thyroid hormone levels from deiodinase loss, muscle weakness from reduced SELENON, increased sun sensitivity, hearing loss, and greater susceptibility to oxidative and cellular stress.16PubMed Central. Human Genetic Disorders Resulting in Systemic Selenoprotein Deficiency A separate gene in the pathway, SEPSECS (the enzyme that performs the final step of selenocysteine biosynthesis on the tRNA), causes a predominantly neurological syndrome with progressive brain atrophy when mutated.16PubMed Central. Human Genetic Disorders Resulting in Systemic Selenoprotein Deficiency
Recent clinical work has expanded the known consequences. Multiple patients with SBP2 mutations have developed aortic root dilation, a dangerous widening of the major blood vessel leaving the heart. In one case, the aorta widened to 8 centimeters by age 23, requiring emergency surgical replacement. In other patients, progressive aortic dilation was tracked over years, with the vessel slowly enlarging from normal to surgically concerning dimensions.17Nature Communications. Selenoprotein deficiency disorder predisposes to aortic aneurysm formation Neurodevelopmental effects have also been described: some children with SBP2 variants present with absent speech, autistic features, and seizures, and while thyroid hormone supplementation improved their motor development, intellectual impairments persisted.18Genetics in Medicine. Expanding the clinical and molecular spectrum of SECISBP2 deficiency
Selenium Deficiency at the Population Level
You do not need a rare genetic mutation to feel the effects of impaired selenoprotein function. Dietary selenium deficiency can do it. The most dramatic historical example is Keshan disease, a dilated cardiomyopathy first identified in northeastern China in regions where the soil is extremely low in selenium. Large-scale epidemiological studies showed that low selenium in local cereal grains and low selenium status in residents tracked closely with the disease. Population-wide supplementation trials using sodium selenite tablets significantly reduced its incidence.19PubMed. An original discovery: selenium deficiency and Keshan disease (an endemic heart disease) Later work showed the picture was not purely nutritional; viral co-infection and genetic susceptibility also played roles, making Keshan disease a gene-environment interaction condition rather than a straightforward deficiency.20PubMed. Is selenium deficiency really the cause of Keshan disease?
The connection between dietary selenium and selenoprotein production is direct. Because selenocysteine is built from selenium supplied by the diet, low selenium intake means fewer charged selenocysteine-tRNAs, less efficient UGA recoding, and reduced selenoprotein output. This is why regions with selenium-poor soil have historically faced higher rates of certain thyroid and cardiovascular conditions. The body does prioritize selenium distribution when supplies are limited, maintaining brain selenoproteins at the expense of liver and muscle, but there are limits to how much compensation is possible.
Why Evolution Kept Selenocysteine Around
Maintaining the selenocysteine system is biologically expensive. It requires dedicated genes for the tRNA, the elongation factor, the SECIS-binding protein, and the biosynthetic enzymes. Given that cysteine can perform many of the same reactions, just more slowly, one reasonable question is why evolution did not simply abandon selenium and stick with sulfur. The answer seems to be that the catalytic advantage is worth the cost, at least under certain conditions.
Comparative genomics tells an interesting story. Among bacteria, roughly one in five sequenced species uses selenocysteine, with the largest selenoproteomes, up to 31 selenoproteins, found in certain groups of anaerobic bacteria. The evolutionary trend appears to favor cysteine-to-selenocysteine replacement: over time, organisms gain selenocysteine in positions that previously held cysteine, presumably because the selenium version works better. Reversals, where selenocysteine is replaced back with cysteine, are much rarer and tend to occur in just a few enzyme families.21PubMed Central. Dynamic evolution of selenocysteine utilization in bacteria: a balance between selenoprotein loss and evolution of selenocysteine from redox active cysteine residues
In animals, the picture is somewhat different. Even primitive animals have a selenoprotein repertoire comparable in range and variety to that of humans. Over evolutionary time, few new selenoprotein families have appeared and few have been lost, with the notable exception of massive, apparently independent losses in nematodes and insects.22PubMed Central. Evolution of selenoproteins in the metazoan Broader analysis across eukaryotes reveals a striking pattern: aquatic organisms tend to have large selenoproteomes, while several lineages of terrestrial organisms have shrunk theirs, either by losing selenoprotein genes outright or by swapping selenocysteine for cysteine. Land plants, fungi, nematodes, and insects have all independently reduced or eliminated their selenoprotein use.23Genome Biology. Evolutionary dynamics of eukaryotic selenoproteomes: large selenoproteomes may associate with aquatic life and small with terrestrial life One hypothesis is that aquatic environments provide more consistent selenium availability, making the investment in selenoprotein machinery more reliably worthwhile.
The 22nd Amino Acid and How It Compares
Selenocysteine is not entirely alone in breaking the 20-amino-acid canon. Pyrrolysine, sometimes called the 22nd amino acid, is also genetically encoded and inserted during translation. It uses the UAG (amber) stop codon instead of UGA. But the two systems are fundamentally different in how they work. Selenocysteine requires the SECIS element, a specialized elongation factor, an on-tRNA biosynthesis pathway, and multiple accessory proteins. Pyrrolysine, by contrast, is made as a free amino acid, loaded directly onto its tRNA by a dedicated enzyme, and inserted at UAG codons without needing any elaborate recoding signals in the mRNA.24PubMed Central. Distinct genetic code expansion strategies for selenocysteine and pyrrolysine are reflected in different aminoacyl-tRNA formation systems
Pyrrolysine is far more restricted in its distribution. Among archaea, it is found almost exclusively in methanogens, the microbes that produce methane. Selenocysteine also shows up in methanogens, and these organisms are the only known archaeal group that uses both non-standard amino acids. Almost all known archaeal pyrrolysine-containing proteins and nearly all confirmed archaeal selenocysteine-containing proteins are involved in the biochemistry of methane production.25PubMed Central. Selenocysteine, pyrrolysine, and the unique energy metabolism of methanogenic archaea In eukaryotes, pyrrolysine has not been found at all, whereas selenocysteine is widespread across animals, many protists, and some algae.
Engineering Selenocysteine Into New Proteins
The catalytic advantages of selenocysteine have attracted the attention of biotechnologists who want to put it into proteins that do not naturally contain it. The challenge is that all the recoding machinery has to be present and functional in whatever production system you are using. Protocols now exist for expressing recombinant selenoproteins in E. coli by using a rewired translation system. These typically involve recoding a UAG stop codon (rather than the native UGA) and co-expressing the necessary selenocysteine insertion factors.26PubMed Central. Introducing Selenocysteine into Recombinant Proteins in Escherichia coli Multiple strategies have been developed to overcome the inherent barriers, including genetic code expansion approaches that borrow elements from different organisms and synthetic biology techniques that redesign the tRNA and elongation factor interactions.27PubMed. Site-Specific Selenocysteine Incorporation into Proteins by Genetic Engineering
Potential applications include designing more active oxidoreductase enzymes for industrial use, creating selenocysteine-containing antibodies with novel reactivity for bioconjugation, and building research tools to probe how redox chemistry works in living cells. The field is still young, but the idea of harnessing the speed and resilience of selenium-based catalysis in engineered proteins is a consistent draw.
Finding Selenoproteins Hidden in Genomes
One practical challenge in selenoprotein biology is simply finding the genes. Because selenocysteine is encoded by UGA, which every gene-finding algorithm recognizes as a stop codon, automated annotation tools routinely truncate selenoprotein genes at the wrong position. A gene that should be annotated as encoding a full-length enzyme gets flagged as ending prematurely at the UGA, and the selenocysteine residue is never recognized. This means selenoprotein genes are chronically under-annotated in genome databases.
Specialized computational tools have been developed to address this. One approach scans bacterial genomes for the characteristic SECIS stem-loop structure downstream of UGA codons, successfully identifying over 96 percent of known selenoprotein genes and predicting new ones.28Bioinformatics. An algorithm for identification of bacterial selenocysteine insertion sequence elements and selenoprotein genes For eukaryotes, profile-based methods use curated alignments of known selenoprotein families to scan genomes for homologous sequences, producing accurate predictions with minimal human intervention.29Bioinformatics. Selenoprofiles: profile-based scanning of eukaryotic genome sequences for selenoprotein genes More recently, deep-learning approaches using transformer-based architectures have shown strong performance in detecting selenocysteine-encoding UGA codons in bacterial genomes, outperforming older methods and identifying previously unknown selenoprotein genes.30PubMed Central. deep-Sep: a deep learning-based method for fast and accurate prediction of selenoprotein genes in bacteria As more genomes are sequenced, these tools are increasingly important for understanding just how widespread and varied selenocysteine use is across the tree of life.