A gene is a stretch of DNA that carries the instructions for building a specific molecule, usually a protein, that your body needs to function. Your cells read these instructions through a two-step copying-and-assembly process, first transcribing the DNA into a messenger molecule and then translating that message into a protein. But that textbook definition, tidy as it sounds, understates how strange and flexible genes really are. Many genes don’t make proteins at all, the same gene can produce different products depending on the tissue or the time of day, and the boundaries of what counts as “a gene” have been shifting for over a century.
How the Meaning of “Gene” Has Changed
When the Danish botanist Wilhelm Johannsen coined the word “gene” in 1909, he meant it as a deliberately vague placeholder for whatever invisible unit was responsible for inherited traits. For the first few decades, scientists treated a gene as an indivisible bead on a string: one unit of inheritance that couldn’t be broken apart, responsible for one function. That classical view held through the 1930s.1PubMed. The concept of the gene: short history and present status
Then researchers discovered that recombination could happen within a gene, not just between genes, which meant genes had internal parts. By the time DNA was confirmed as the physical carrier of inheritance, the picture became more concrete: a gene was a segment of DNA whose subunits were individual nucleotides, and each gene was believed to be responsible for producing one messenger RNA and, through it, one protein.2The Journal of Medicine and Philosophy: A Forum for Bioethics and Philosophy of Medicine. Historical Development of the Concept of the Gene That neat “one gene, one protein” idea survived into the 1970s before molecular discoveries started dismantling it. Overlapping genes, genes that produce multiple different proteins, genes that don’t make proteins at all: none of these fit the old framework. The definition keeps evolving because genes keep turning out to be more complicated than the last generation of scientists assumed.3Genetics. The Evolving Definition of the Term “Gene”
What a Gene Looks Like in DNA
At the molecular level, a gene is a specific sequence of the four chemical bases (A, T, G, and C) strung along a DNA molecule. In humans and other complex organisms, a gene’s coding sequence is typically broken into segments called exons, separated by non-coding stretches called introns. A study cataloging the structure of genes across several model organisms found that, on average, a gene contains roughly four introns per thousand base pairs of protein-coding sequence, with most exons encoding only about 30 to 40 amino-acid residues and most introns being relatively short.4Oxford Academic (Nucleic Acids Research). Intron—exon structures of eukaryotic model organisms The introns get removed later, during processing. This patchwork layout is one of the reasons a single gene can yield different proteins, because the cell can stitch the exons together in different combinations.
This gene structure is not universal. Bacteria generally lack introns, so their genes are continuous stretches of coding DNA, often clustered into groups called operons where several genes sit next to each other and get read as a single unit.5PubMed Central. Noncontiguous operon is a genetic organization for coordinating bacterial gene expression This arrangement lets bacteria coordinate the production of related proteins efficiently. Eukaryotic genes, by contrast, are typically read one at a time and regulated independently, even when they sit near each other on the chromosome.
From Gene to Protein
The process of turning a gene’s DNA sequence into a working protein happens in two main stages. The first, transcription, takes place in the cell’s nucleus. An enzyme called RNA polymerase II reads along the DNA and builds a complementary strand of messenger RNA (mRNA). This isn’t a passive copying job. RNA polymerase II recruits a variety of helper proteins throughout the process, and a special tail on the enzyme coordinates the processing steps that happen while the mRNA is being made, including capping the front end and adding a protective tail to the back end.6PubMed. RNA polymerase II conducts a symphony of pre-mRNA processing activities
Before the mRNA leaves the nucleus, the introns have to be cut out and the exons spliced together. This is where things get interesting, because the cell doesn’t always splice the same way. A process called alternative splicing allows different combinations of exons to be joined, producing different protein variants from the same gene. This is a major source of protein diversity and plays a role in cell specialization and development.7PubMed Central. Mechanism of alternative splicing and its regulation A single human gene can, through alternative splicing, give rise to dozens of distinct proteins. This is one reason humans have roughly 20,000 protein-coding genes but many more distinct proteins.
The second stage, translation, happens outside the nucleus on molecular machines called ribosomes. The ribosome reads the mRNA three bases at a time, and each three-base “codon” specifies a particular amino acid. Small adaptor molecules called transfer RNAs carry the correct amino acid to the ribosome, where they line up in the sequence dictated by the mRNA. The amino acids get linked together into a chain, which then folds into a functional protein.8Reports on Progress in Physics. The ribosome and the mechanism of protein synthesis
What Happens After the Protein Is Built
Translation is not the final step. Once a protein chain is assembled, it often undergoes chemical modifications that alter its shape, location, or activity. These post-translational modifications are found across all forms of life and give proteins a much wider range of functions than the gene sequence alone would suggest.9PubMed Central. Catalytic activity regulation through post-translational modification: the expanding universe of protein diversity Common modifications include adding small chemical groups like phosphate or sugar molecules to specific amino acids, clipping off sections of the protein, or attaching it to a fatty acid that anchors it in a cell membrane.
This matters for understanding what genes actually “do” because the relationship between a gene and its final product is not as direct as the DNA-to-protein pipeline might suggest. The same protein, produced from the same gene, can behave differently depending on which modifications it receives, and those modifications respond to signals from the cell’s environment. A gene provides the blueprint, but the finished product reflects layers of editing, remodeling, and context.
Not All Genes Make Proteins
One of the biggest surprises of modern genomics is that only a small fraction of the human genome codes for proteins, yet much of the rest is far from useless. Non-coding RNAs, molecules transcribed from DNA but never translated into protein, make up the majority of the genome’s output. These molecules serve a wide range of roles: some regulate whether other genes get turned on or off, some help assemble the protein-building machinery, and some act as molecular sponges that soak up other regulatory RNAs to dampen their effects.10PubMed Central. MicroRNAs (miRNAs) and Long Non-Coding RNAs (lncRNAs) as New Tools for Cancer Therapy: First Steps from Bench to Bedside They participate in processes as varied as immune response, embryonic development, and tumor suppression.
The ENCODE project, a massive effort to catalog functional elements across the human genome, assigned biochemical functions to about 80 percent of the genome, with much of that activity occurring outside protein-coding regions.11PubMed Central. An integrated encyclopedia of DNA elements in the human genome There is ongoing scientific debate about how much of that activity is truly functional versus biochemical noise, but the old idea that non-coding DNA is “junk” has been thoroughly challenged. When someone asks “what is a gene,” the honest modern answer has to include genes whose product is an RNA molecule, not a protein.
How Genes Get Turned On and Off
Every cell in your body carries essentially the same complete set of genes, yet a liver cell behaves nothing like a brain cell. The difference comes down to gene regulation: which genes are active, how strongly they’re active, and when. Several layers of control determine this.
One important layer involves regions of DNA called enhancers. These are stretches of sequence, sometimes located far away from the gene they regulate, that boost a gene’s transcription when activated. Enhancers are central to making sure genes fire in the right cells at the right time during development and in everyday tissue maintenance.12PubMed Central. What is an enhancer? They work by physically looping through three-dimensional space to make contact with the gene’s starting point, sometimes bridging tens of thousands of bases of intervening DNA.
This looping doesn’t happen randomly. The genome folds into neighborhoods called topologically associating domains, or TADs, which keep enhancers and their target genes in the same physical compartment and prevent enhancers from accidentally activating genes in neighboring compartments.13PubMed Central. Topologically Associating Domains and Regulatory Landscapes in Development, Evolution and Disease A protein called cohesin helps form these loops, and when cohesin is lost, TADs can split apart and gene activity changes in context-dependent ways.14PubMed Central. Context-dependent perturbations in chromatin folding and the transcriptome by cohesin and related factors Think of it as the genome being organized into filing cabinets: the physical folding determines which regulatory switches are close enough to reach which genes.
Epigenetic Marks and Heritable Changes Without DNA Mutations
Beyond the sequence itself, genes can be silenced or activated by chemical marks added on top of the DNA or on the proteins that package it. These epigenetic modifications don’t change the underlying genetic code, but they change whether a gene is accessible to the transcription machinery. The two best-studied types are DNA methylation, where a small chemical group is attached directly to certain bases, and histone modifications, where the spool-like proteins that DNA wraps around get tagged with various chemical groups. Together, these marks regulate how tightly or loosely a stretch of DNA is packed, and therefore how easily it can be read.15PubMed Central. Epigenetic modifications: basic mechanisms and role in cardiovascular disease
What makes epigenetics particularly interesting is that some of these marks can be passed along when a cell divides, and in certain cases across generations. They’re also sensitive to environmental influences like diet, stress, and chemical exposures. This means the question “how do genes work” can’t be answered by looking at DNA sequence alone. The packaging and marking of that DNA are an essential part of the story.
Genetic Variation Between People
No two people (except identical twins) share exactly the same DNA sequence. The simplest and most common form of variation is the single nucleotide polymorphism, or SNP, where one base differs between individuals at a particular spot in the genome. SNPs occur roughly once every thousand base pairs and account for much of the diversity you can see between people, from hair texture to drug metabolism. Some SNPs change the amino acid a gene codes for. Others are “silent” in that they don’t change the protein, but even silent SNPs can affect how much protein gets made or how stable the mRNA is.16PubMed. SNPs: impact on gene function and phenotype
Larger-scale variation exists too. Entire genes or segments of genes can be duplicated or deleted. These copy-number variations are one of the fundamental mechanisms behind the evolution of new gene functions: when a gene gets duplicated, one copy can keep doing the original job while the other accumulates mutations and potentially takes on a new role. Many duplicated copies fail and become non-functional pseudogenes, but occasionally positive selection preserves a beneficial new variant, and that’s how gene families with new capabilities arise over evolutionary time.17PubMed Central. The current excitement about copy-number variation: how it relates to gene duplications and protein families18PubMed Central. Simulating evolution by gene duplication
Why Most Traits Don’t Follow Simple Inheritance Rules
Introductory biology often teaches genetics through examples like Mendel’s peas, where a single gene with two versions produces a clear either/or trait. In reality, most human characteristics, and most common diseases, don’t work this way. Conditions like diabetes, heart disease, and psychiatric disorders are polygenic: they involve many genes, each contributing a small individual effect, with the total risk depending on the combined burden of all those small contributions plus environmental factors.19PubMed Central. Discovery and implications of polygenicity of common diseases
Researchers now use polygenic risk scores, which tally up the effects of many risk-associated SNPs across the genome, to estimate an individual’s genetic susceptibility to a given condition. These scores can identify people at elevated risk, potentially allowing earlier screening or prevention. But because each individual genetic variant has such a tiny effect, no single gene test can predict most common diseases with certainty. The “gene for X” framing that shows up in headlines is almost always an oversimplification. For complex traits, it’s more accurate to think of hundreds or thousands of genes nudging a trait in one direction, with lifestyle and environment pushing alongside them.
Genes Respond to the Environment
Genes don’t operate in a vacuum. Environmental signals directly influence which genes are active and how strongly they’re expressed. This is vividly demonstrated in plants, which can’t relocate when conditions change and instead must reprogram their gene expression on the fly. Plants have evolved systems where the same molecular pathways that detect light also integrate signals about temperature and water availability, allowing a single network to coordinate the organism’s response to multiple changing conditions at once.20PubMed Central. Light signaling as cellular integrator of multiple environmental cues in plants21PubMed Central. Circadian and environmental signal integration in a natural population of Arabidopsis
In humans, gene-environment interactions are equally important, though harder to study. Nutrition, physical activity, toxin exposure, sleep, and psychological stress all alter gene expression patterns in various tissues. This is part of why identical twins, who share the same DNA, can develop different diseases over a lifetime. Their genes are the same; the environments those genes responded to were not.
Genes Outside the Nucleus
Not all of your genes live on your chromosomes. Mitochondria, the energy-producing structures inside your cells, carry their own small circular genome with 37 genes. Mitochondrial DNA (mtDNA) follows a distinctive inheritance pattern: it comes exclusively from your mother.22PubMed. Maternal inheritance of mitochondrial DNA by diverse mechanisms to eliminate paternal mitochondrial DNA Sperm do deliver mitochondria into the egg at fertilization, but paternal mtDNA is eliminated.
Recent research has clarified the mechanism behind this. During sperm development, cells produce a version of a key mitochondrial protein (TFAM) that gets rerouted to the sperm cell’s nucleus instead of entering its mitochondria. Without TFAM to protect and maintain the mitochondrial DNA, the mtDNA in mature sperm is degraded before the sperm ever reaches the egg.23Nature Genetics. Molecular basis for maternal inheritance of human mitochondrial DNA The result is that mitochondrial genes pass mother to child across generations, making mtDNA useful for tracing maternal lineages and also clinically relevant: mutations in mitochondrial genes cause a distinct set of inherited diseases that follow a maternal pattern.
Overlapping Genes and Genomic Complexity
One reason the gene concept keeps getting messier is that genes don’t always occupy their own private stretch of DNA. In the human genome, researchers have found hundreds of cases where two genes overlap, meaning they’re transcribed from the same piece of DNA but in opposite directions. An analysis estimated a minimum of about 900 such overlapping transcript pairs across the human genome, with the overlaps typically spanning 50 to 200 nucleotides, sometimes extending into the protein-coding region of one of the transcripts.24PubMed Central. Overlapping Antisense Transcription in the Human Genome These overlaps mean that a mutation at a single position could affect two different genes simultaneously, and that the regulation of one gene can influence its neighbor through the shared DNA. The tidy image of genes as discrete, well-separated units along a chromosome is more of a convenient abstraction than an accurate description of what’s actually going on.
Editing Genes with CRISPR
Understanding how genes work has opened the door to deliberately changing them. The most prominent tool for this is CRISPR-Cas, a system originally discovered as a defense mechanism in bacteria that has been repurposed into a precise gene-editing technology.25PubMed. CRISPR/Cas in the era of nanomedicine and synthetic biology The system uses a short guide RNA to direct an enzyme to a specific location in the genome, where it cuts the DNA. The cell’s repair machinery then fixes the break, and researchers can exploit that repair process to delete, correct, or insert genetic sequences.
Recent advances have pushed the precision of CRISPR to the level of editing single nucleotides, making it possible in principle to correct the kind of point mutations that cause many genetic diseases.26PubMed Central. Recent Advances in CRISPR-Cas Technologies for Synthetic Biology Clinical trials are now underway for conditions ranging from sickle cell disease to certain inherited forms of blindness. The therapeutic promise is real, but so are the challenges: delivering the editing machinery to the right cells, avoiding unintended cuts elsewhere in the genome, and navigating the ethical questions around editing human embryos. Gene editing rests on everything described in this article, from understanding gene structure and regulation to knowing how variation arises and what happens when a sequence changes. It’s the practical culmination of a century of asking what a gene is and how it works.