De Novo Genes: Impact, Discovery, and Future Insights

De novo genes are genes that arise from stretches of DNA that previously did not encode any protein or functional RNA. Once considered an impossibility in mainstream molecular evolution, their existence is now firmly established across species from yeast and fruit flies to plants and humans. Comparative genomics and transcriptomics have demonstrated that new genes do not always emerge from copies of old ones; sometimes they appear, essentially, from scratch. That shift in understanding has opened questions about how these genes affect organisms, how scientists find them, and what their discovery means for everything from brain evolution to cancer treatment.

From Non-Coding DNA to Working Gene

For decades, the standard model of gene evolution held that new genes could only come from preexisting genes, mainly through a process where a gene gets duplicated and one copy gradually changes function. De novo gene birth upends that assumption. A stretch of DNA that was doing nothing recognizable as “gene-like” accumulates mutations over evolutionary time until it contains the right signals to be transcribed, translated, or both.1PLoS Genetics. De novo gene birth The result is a completely new protein or functional RNA with no ancestor in any other lineage.

What makes this even more surprising is how ordinary the raw material looks. The sequence properties of de novo proteins are nearly indistinguishable from randomly generated sequences, yet these proteins can gain functions and integrate into existing cellular networks with relative ease.2ScienceDirect (Current Opinion in Structural Biology). Structure and function of naturally evolved de novo proteins That finding has forced researchers to rethink what it takes for a protein to be biologically useful. You do not need an elaborate, deeply conserved fold to participate in cell biology; something cobbled together from formerly junk DNA can sometimes do the job.

One concrete example illustrates the process at the nucleotide level. A study comparing human and mouse genomes found that a human gene containing a 107-amino-acid open reading frame has no equivalent protein-coding sequence in the mouse. The homologous mouse region is non-coding, containing only two start codons followed by tiny peptides with no match to the human protein or any known protein in other organisms.3bioRxiv. Evolutionary formation of a human de novo open reading frame from a mouse non-coding DNA sequence via biased random mutations The human version arose through accumulated mutations that, step by step, built a functional reading frame where none existed before.

How Scientists Find De Novo Genes

Identifying a de novo gene is harder than it sounds. The challenge is distinguishing a genuinely new gene from one that simply evolved so fast its ancestor is no longer recognizable by standard sequence-search tools. If two related species each have a gene in the same chromosomal neighborhood but the sequences have diverged beyond recognition, a naive search would incorrectly call each one “new.” Synteny-based methods address this problem by checking whether genes sit in conserved positions across genomes, allowing researchers to identify candidate homologs even when sequence similarity has been erased.4PubMed Central. Synteny-based analyses indicate that sequence divergence is not the main source of orphan genes

A complementary approach called phylostratigraphy traces how genes distribute across the tree of life. By mapping when each gene first appears in lineage history, researchers can estimate the age of genes and look for bursts of new gene emergence at particular evolutionary nodes. One such analysis used the well-annotated mouse genome as a reference and found patterns consistent with frequent de novo gene evolution, rather than the rare-event scenario older models predicted.5PubMed Central. Phylogenetic patterns of emergence of new genes support a model of frequent de novo evolution

Neither approach alone is definitive. Phylostratigraphy can overcount de novo genes if fast-diverging sequences are mistaken for novel ones, and synteny methods can miss genes that arose in genomic regions with poor conservation of chromosome structure. The field increasingly combines both approaches, along with direct experimental evidence of translation, to build a more reliable census.

Catching Translation in the Act

One of the most powerful experimental tools for confirming that a stretch of DNA actually produces protein is ribosome profiling, a technique that captures snapshots of ribosomes actively translating mRNA. Ribosome profiling studies have revealed that translation is far more pervasive than gene annotations suggest. Ribosomes occupy many regions of the transcriptome previously thought to be non-coding, including long non-coding RNAs and untranslated regions of messenger RNAs. These footprints show hallmarks of genuine translation: they associate with the large ribosomal subunit, respond to drugs that target the translation machinery, and display the characteristic three-nucleotide reading pattern.6PubMed Central. Ribosome profiling reveals pervasive translation outside of annotated protein-coding genes

Not every ribosome footprint on a transcript means a functional protein is being made, though. Researchers have developed scoring systems to separate truly protein-coding regions from non-coding ones. One such metric, the ribosome release score, measures whether ribosomes drop off at a stop codon the way they do for known genes. For established protein-coding genes, ribosome occupancy in the coding region is roughly 112-fold higher than in the region after the stop codon. For long non-coding RNAs, by contrast, the ratio hovers around 1, indicating roughly equal occupancy before and after the stop codon, a hallmark of non-productive ribosome scanning rather than real translation.7PubMed Central. Ribosome profiling provides evidence that large non-coding RNAs do not encode proteins Tools like this help researchers sort through the noise and zero in on genuinely novel translated products.

The approach works across kingdoms. In Arabidopsis, a high-resolution version of ribosome profiling uncovered small open reading frames hiding in annotated non-coding RNAs and pseudogenes, regions assumed to be evolutionary dead ends.8PubMed Central. Super-resolution ribosome profiling reveals unannotated translation events in Arabidopsis The picture that emerges is one of a genome with far more translational activity than its formal annotation acknowledges, providing fertile ground for de novo gene birth.

The Testis as an Evolutionary Nursery

If de novo genes can pop up anywhere, why do they keep showing up in the same place? Across species, new genes are disproportionately expressed in the testes. A single-cell RNA sequencing study of fruit fly testes found that lineage-specific de novo genes are commonly expressed in early spermatocytes, and spermatocytes show the highest relative abundance of both fixed and still-segregating de novo genes compared to other cell types.9PubMed Central. Testis single-cell RNA-seq reveals the dynamics of de novo gene transcription and germline mutational bias in Drosophila Population-level sampling in Drosophila melanogaster identified 142 segregating and 106 fixed testis-expressed de novo genes, giving a sense of the sheer volume of novel genetic material passing through this organ.10PubMed Central. Origin and spread of de novo genes in Drosophila melanogaster populations

The reasons are partly mechanical. The testes have unusually permissive gene regulation: chromatin is more open, and transcription is less tightly controlled than in most other tissues. That permissiveness means random open reading frames in non-coding DNA are more likely to get transcribed there first. Once a proto-gene is being expressed in the testes, natural selection can act on it if the resulting protein or peptide has any effect on reproductive success. Over evolutionary time, some of these testis-expressed newcomers get recruited into expression in other tissues, effectively “graduating” from their initial nursery.

De Novo Genes in Human Brain Evolution

Some of the most striking examples of de novo gene function involve the human brain. A human-specific de novo gene called SP0535 is preferentially expressed in the ventricular zone of the developing fetal brain. When researchers knocked it out in human stem-cell-derived brain organoids, growth and neuron production were compromised. In mice engineered to carry the gene, SP0535 drove expansion of the fetal cortex and even induced the formation of fold-like structures resembling the ridges and grooves of the human brain. The transgenic mice also showed improved cognitive ability and working memory.11PubMed Central. A Human-Specific De Novo Gene Promotes Cortical Expansion and Folding SP0535 is the first demonstrated case of a human-specific de novo gene that promotes cortical folding, and it does so by slotting into an existing conserved molecular network rather than building a new one.

Another candidate, FLJ33706, was identified through a computational screen of genetic factors linked to nicotine addiction and brain function. Its mRNA and protein are most abundant in the brain, and immunohistochemistry localizes the protein to neurons in the cortex, cerebellum, and midbrain.12PLoS Computational Biology. A human-specific de novo protein-coding gene associated with human brain functions That a gene born from non-coding DNA can end up playing a role in something as complex as cortical folding or neuronal signaling underscores how consequential de novo gene birth can be for a lineage’s evolutionary trajectory.

De Novo Genes and Cancer

The same features that make testes and brain tissue hospitable to de novo gene expression also make tumors a place where these genes can become active in dangerous ways. Because tumors often loosen the normal controls on gene expression, de novo genes that are silent or barely expressed in healthy tissue can become highly upregulated in cancerous cells. An analysis of over 5,200 tumor samples across 22 cancer types found that young human de novo genes showed significant upregulation in tumors. About two-thirds of the upregulated set had no detectable expression whatsoever in matched normal tissues, meaning their expression was entirely new in the tumor context.13Cell Genomics. Oncogenic roles of young human de novo genes and their potential as neoantigens in cancer immunotherapy This expansion of expression was far more prevalent among de novo genes than among background genes, and it occurred across a broader spectrum of tumor types.

One well-characterized example is PBOV1, a human de novo gene with tumor-specific expression. Researchers detected PBOV1 in tumors across 16 different tissue types, including brain, lung, liver, colon, ovary, and prostate, and confirmed expression in 22 out of 31 clinical tumor samples from an independent panel.14PLOS ONE. PBOV1 Is a Human De Novo Gene with Tumor-Specific Expression That Is Associated with a Positive Clinical Outcome of Cancer Interestingly, PBOV1 expression was associated with a positive clinical outcome, complicating any simple narrative that tumor-expressed de novo genes are uniformly harmful.

The broader picture is tantalizing for immunotherapy. Because many of these de novo gene products are absent from normal tissues, the immune system has never learned to tolerate them. That makes them potential neoantigens, proteins that the immune system could be trained to attack specifically in tumors without harming healthy cells. Research in this area is still early, but the sheer number of cancer types showing de novo gene upregulation suggests a wide potential target landscape.

Viral Overprinting as a Parallel Route

Viruses face extreme pressure to pack maximum function into minimal genome space, and they have evolved their own version of de novo gene birth through a process called overprinting. Mutations within an existing coding sequence can create a new protein in a different reading frame, overlapping the original gene. In unspliced RNA viruses infecting eukaryotes, researchers confirmed that 17 overlapping proteins had been created de novo across 43 viral genera.15PubMed Central. Overlapping genes produce proteins with unusual sequence properties and offer insight into de novo protein creation

More recent work using discriminant analysis of codon usage has refined the ability to distinguish ancestral reading frames from de novo ones in overlapping genes. For genes read in the +1 frame relative to the ancestral sequence, the method achieved 97.2% classification accuracy; for the +2 frame, accuracy reached 100%.16Virology. New insights into the evolutionary features of viral overlapping genes by discriminant analysis These viral de novo proteins share certain properties with the de novo genes found in cellular organisms: their sequences look unusual compared to well-established genes, and they have no detectable homologs outside their lineage. Viruses, with their rapid mutation rates and enormous population sizes, provide a kind of time-lapse view of how novel proteins can be born from previously coding sequences, paralleling the cellular route from non-coding DNA.

How New Genes Wire Into Existing Networks

A de novo gene that produces a protein is not automatically useful. To have a biological impact, its product needs to participate in the molecular networks that run the cell. One way this happens is through enhancers, the regulatory elements that control when and where genes are turned on. A comparative study in mice found that young, species-specific open reading frames in non-coding regions are preferentially located near enhancers, while older open reading frames are not.17Molecular Biology and Evolution. Enhancers Facilitate the Birth of De Novo Genes and Gene Integration into Regulatory Networks Proximity to enhancers gives these proto-genes a ready-made on switch, allowing them to be transcribed without first evolving their own promoter from scratch.

Over evolutionary time, de novo genes accumulate more regulatory connections. The same study found a positive correlation between a gene’s age and its number of enhancer interactions: mouse-specific open reading frames had a median of about 10 enhancer interactions, while those shared with more distantly related species had medians of 13, 15, and 21 interactions, scaling with evolutionary age.17Molecular Biology and Evolution. Enhancers Facilitate the Birth of De Novo Genes and Gene Integration into Regulatory Networks Young genes located inside existing genes had more enhancer connections from the start, likely because they were borrowing the regulatory apparatus of their host gene. This piggyback mechanism may be one of the main ways de novo genes get their foot in the regulatory door.

The evidence of functional integration goes beyond regulation. In yeast, the de novo gene BSC4 shows signs of purifying selection, meaning mutations that break it are weeded out by natural selection. It is expressed at both the mRNA and protein levels, and deleting it is lethal when combined with the loss of two other yeast genes.18PLoS Genetics. De novo gene birth – Section: 3 Prevalence of de novo gene birth A gene that is under purifying selection and participates in genetic interactions is no longer a bystander; it has become part of the cellular infrastructure.

Stress Adaptation in Plants

De novo genes are not limited to animal nervous systems and tumors. In Arabidopsis, a de novo gene called SWK enhances seed germination under osmotic stress, the kind of stress plants face during drought. Researchers found that natural populations carrying SWK tend to live in areas with significantly lower rainfall in the driest quarter compared to populations lacking the gene, at least in African accessions.19Molecular Biology and Evolution. A de novo Gene Promotes Seed Germination Under Drought Stress in Arabidopsis The implication is that a gene born from non-coding DNA may have improved the plant’s ability to colonize drought-prone habitats, providing a concrete example of de novo gene birth driving ecological adaptation rather than just molecular novelty.

Plant genomes, with their frequent whole-genome duplications and large non-coding fractions, may be especially hospitable to de novo gene emergence. But plants have been less studied in this context than animals and yeast, partly because the computational pipelines were originally built around animal genomes. That is starting to change with the development of machine-learning approaches tailored to plant genomes.

Machine Learning and the Future of Detection

Identifying de novo genes has traditionally required laborious manual curation: aligning genomes, checking synteny, running phylostratigraphic analyses, and then validating candidates experimentally. Machine-learning algorithms are beginning to speed up the process. A study using decision tree and neural network models trained on known de novo genes from three flowering plant species found that both algorithms achieved high accuracy and recall when tested within genomes, and could even predict de novo genes across species.20bioRxiv. Accurate identification of de novo genes in plant genomes using machine learning algorithms If these tools prove robust across broader taxonomic ranges, they could dramatically expand the catalog of known de novo genes, particularly in organisms where genomic resources are thin.

The practical payoff of better detection extends to medicine. If de novo genes are reliably upregulated in tumors and absent from healthy tissue, they become attractive immunotherapy targets, but only if you can find them systematically. Machine-learning pipelines that flag candidate de novo genes in newly sequenced genomes could feed directly into neoantigen prediction workflows, turning evolutionary curiosity into clinical utility. The field is still building its foundational tools, but the trajectory points toward de novo genes becoming a routine component of genome annotation rather than an exotic afterthought.

Why Random-Looking Sequences Can Still Work

One of the persistent puzzles about de novo genes is that their protein products look, by standard metrics, like random amino acid chains. They typically lack the conserved domains, recognizable folds, and sequence signatures that characterize well-established proteins. Research into synthetic random-sequence libraries has confirmed that unevolved polypeptides can fold into soluble structures and, in some cases, acquire specific biochemical functions through in vitro evolution.21PubMed Central. De novo proteins from random sequences through in vitro evolution This suggests that the protein universe is far more permissive than was once assumed: the space of sequences that can do something useful is much larger than the space currently occupied by known protein families.

That permissiveness has implications beyond evolutionary biology. If random-looking sequences can be functional, then protein engineering efforts do not need to start from known scaffolds. De novo protein design, an active area in biotech, draws on the same principle: you can sometimes build useful molecular machinery without borrowing from nature’s existing toolkit. The de novo genes that evolution stumbles upon and the de novo proteins that engineers design are, in a sense, exploring the same vast and surprisingly hospitable sequence space from different directions.

Leave a Reply

Your email address will not be published. Required fields are marked *