How Many Human Proteins Are There?

The human genome contains roughly 20,000 protein-coding genes, but the number of distinct proteins those genes produce is vastly larger and, honestly, still unknown. Depending on how you define “protein,” credible estimates range from just under 20,000 (one per gene, ignoring everything else) to potentially millions of distinct molecular forms circulating through your cells at any given moment. The gap between those numbers is where most of the interesting biology lives.

The Gene Count Took Decades to Settle

Before anyone could count human proteins, they needed to count the genes that encode them. That turned out to be surprisingly difficult. In the 1960s, early calculations based on the total amount of DNA in human sperm cells and the assumed length of a typical gene suggested the number could be enormous.

When the draft human genome sequence was published in 2001, the initial tally of protein-coding genes came in around 25,000, far lower than the 100,000 many researchers had predicted. But rather than settling the question, the genome sequence kicked off years of revisions. Gene count estimates bounced between 30,000 and 70,000 in the years that followed, as annotation methods improved and researchers debated what qualified as a real gene versus genomic noise.1PubMed. Has the yo-yo stopped? An assessment of human protein-coding gene number The question “how many genes?” was widely expected to be answered once the genome was published, but estimates kept fluctuating.2PubMed Central. Between a chicken and a grape: estimating the number of human genes

Today, the three major gene annotation databases agree on around 19,400 to 21,900 protein-coding genes, though they still disagree on the status of roughly 2,600 of them. That is about one in eight annotated genes where experts cannot agree on whether a sequence genuinely codes for a functional protein.3PubMed Central. The state of the human coding gene catalogues So even the baseline number, the count of genes before you think about anything else, still has meaningful uncertainty baked in.

Why One Gene Does Not Mean One Protein

If each gene produced exactly one protein, counting proteins would be simple. But that is not how it works. A single gene can give rise to multiple distinct protein products through a process called alternative splicing, in which sections of a gene’s initial transcript are included or excluded in different combinations before the final messenger RNA is translated. Most human genes undergo alternative splicing, and the result is a library of related but structurally different proteins from one genetic blueprint.

How different are these variants from each other? Research profiling hundreds of isoform pairs found that the majority share fewer than half of their protein interaction partners. In the broader network of cellular interactions, alternative isoforms often behave more like entirely distinct proteins than like minor tweaks of each other.4PubMed Central. Widespread Expansion of Protein Interaction Capabilities by Alternative Splicing The interaction partners unique to specific isoforms tend to show up in particular tissues and belong to distinct functional groups, meaning these variants are not just molecular noise but have tissue-specific jobs.

That said, there is a long-running debate about how many of these splice variants actually matter. Research from the ENCODE project argued that for the vast majority of alternatively spliced isoforms, there is little evidence they function as conventional proteins, and it seems unlikely that the range of standard protein functions can be substantially expanded through splicing alone.5PubMed Central. The implications of alternative splicing in the ENCODE protein complement Work integrating RNA sequencing with mass spectrometry has shown that while changes in transcript usage do alter protein abundance in proportion to transcript levels, the relationship between the transcriptome and the proteome is not one-to-one.6PubMed Central. Impact of Alternative Splicing on the Human Proteome Many transcripts either never get translated or produce proteins too unstable to accumulate. So while splicing theoretically generates tens of thousands of variant transcripts, the number that generate stable, functional proteins in your cells is almost certainly much smaller than the RNA data alone suggests.

Proteoforms and the Chemical Modification Explosion

Even after a protein is translated, it is not finished. Cells attach chemical groups, sugars, lipids, and even small proteins to specific spots on the chain. These post-translational modifications, or PTMs, can change how a protein folds, what it binds to, where it goes in the cell, and whether it is active or dormant. PTMs are not rare edge cases: they are a core mechanism by which cells regulate protein function and massively expand the diversity of the proteome.7PubMed Central. Proteome-wide profiling and mapping of post translational modifications in human hearts A single protein can carry multiple types of modifications at multiple sites, and different combinations produce distinct molecular species.

This is where the concept of the “proteoform” comes in. A proteoform is a single specific molecular form of a protein, defined by its exact amino acid sequence (including any genetic variation and splice variation) plus the precise set of chemical modifications it carries. The term was coined to capture the enormous combinatorial space that opens up when you layer genetic variants, splice variants, and PTMs on top of one another. One analysis framing the question asked how many distinct proteoforms arise from roughly 20,300 human genes, and the theoretical space is staggering.8PubMed Central. How many human proteoforms are there? The Human Proteoform Project has described proteoforms as the products of combinations of genetic polymorphisms, RNA splice variants, and post-translational modifications.9PubMed Central. The Human Proteoform Project: Defining the human proteome

Estimates of the total number of human proteoforms vary wildly. Conservative counts that stick to well-documented modifications on well-characterized proteins land in the hundreds of thousands. More expansive calculations that include all theoretically possible modification combinations push into the millions or even billions. The honest answer is that nobody has measured the full proteoform landscape, and current technology cannot do so comprehensively.

Proteins Hiding in Non-Coding Regions

For decades, genome annotation pipelines used a size cutoff: if a stretch of DNA was too short to encode a protein of at least 100 amino acids, it was ignored. That assumption turns out to have been wrong. Small open reading frames scattered across the genome, including in regions classified as “non-coding,” can and do produce tiny proteins called microproteins. These microproteins interact with other molecules to regulate gene expression and participate in biological processes.10PubMed Central. Small open reading frame-encoded microproteins in cancer: identification, biological functions and clinical significance

Early genome annotation dismissed these small open reading frames as meaningless noise. It was only with the advent of ribosome profiling, a technique that captures which RNA sequences are actively being translated, that researchers realized many of these “non-coding” transcripts were in fact being read by ribosomes.11PubMed Central. Short open reading frames (sORFs) and microproteins: an update on their identification and validation measures Whether a translated sequence produces a stable, functional microprotein requires further validation, but the catalog of confirmed cases is growing steadily. Researchers combining ribosome profiling with mass spectrometry have identified over a hundred unannotated microproteins in single experiments.12Nature Communications. Rp3: Ribosome profiling-assisted proteogenomics improves coverage and confidence during microprotein discovery

Closely related to microproteins is a broader class of non-canonical proteins produced by what researchers call “cryptic translation.” These come from regions of the genome not traditionally expected to produce proteins: the untranslated regions of messenger RNAs, long non-coding RNAs, pseudogenes, and even frameshifted versions of known genes. One study of three human B-cell lymphomas identified over 14,000 proteins, of which about 2,500 were non-canonical. Of those, roughly 72% were cryptic proteins encoded by ostensibly non-coding regions or frameshifted genes.13PubMed Central. Most non-canonical proteins uniquely populate the proteome or immunopeptidome If these findings generalize, the true protein complement of human cells is substantially larger than what standard gene catalogs predict.

RNA Editing Adds Yet Another Layer

Before messenger RNA is even translated, enzymes called ADARs can chemically modify individual nucleotides in the transcript, converting adenosine to inosine. Because the cell’s translation machinery reads inosine as guanosine, this swap can change which amino acid gets inserted into the protein, producing a protein that differs from what the DNA sequence would predict.14PubMed. Proteome diversification by adenosine to inosine RNA editing This means that even two cells carrying identical genomes and splicing the same transcript can produce different protein products depending on how aggressively the editing enzymes have acted on that particular RNA molecule.

In humans, the number of coding-region editing events that meaningfully change protein sequences is relatively modest compared to the massive amount of editing that occurs in non-coding regions. But each editing event that does hit a coding site creates a protein variant that exists nowhere in the genome. The functional consequences can be dramatic: in the nervous system, RNA editing of ion channel and receptor transcripts fine-tunes signaling properties in ways that have clear physiological effects.15RNA Editing. Adenosine-to-inosine RNA editing: substrates and consequences So while RNA editing does not add thousands upon thousands of new proteins, it adds an unpredictable, environmentally responsive layer of diversity on top of everything else.

How Far Along Is the Search

The HUPO Human Proteome Project has been systematically working to confirm the existence of every predicted human protein using direct experimental evidence such as mass spectrometry and antibody detection. Progress has been steady but slow. In 2013, only about 13,700 proteins had strong experimental support.16PubMed Central. Metrics for the Human Proteome Project 2013-2014 and strategies for finding missing proteins By 2020, that number had climbed past 17,800, representing over 90% of the roughly 19,800 predicted coding genes at that time.17PubMed Central. Research on the Human Proteome Reaches a Major Milestone: >90% of Predicted Human Proteins Now Credibly Detected, According to the HUPO Human Proteome Project

As of 2024, the tally stands at about 18,100 proteins with strong evidence out of roughly 19,400 predicted protein-coding genes, pushing the detection rate to 93%. That still leaves around 1,270 “missing proteins” with no or inadequate experimental confirmation.18PubMed Central. The 2024 Report on the Human Proteome from the HUPO Human Proteome Project These missing proteins tend to be the ones expressed at very low levels, in very specific tissues, or only under unusual conditions, making them genuinely hard to find.

It is worth noting that this 93% figure refers only to the “one protein per gene” baseline. It says nothing about how many splice variants, proteoforms, or microproteins have been characterized. The baseline is close to complete; the full diversity picture is nowhere near it.

Why the Full Proteome Is So Hard to Measure

A major obstacle is dynamic range. In blood plasma alone, protein concentrations span more than ten orders of magnitude. Albumin, the most abundant plasma protein, accounts for roughly half of all protein by weight, and the top 22 proteins account for 99%.19PubMed Central. Mass Spectrometry-Based Plasma Proteomics: Considerations from Sample Collection to Achieving Translational Data That means the rarest proteins are drowned out by a handful of extremely abundant ones. No current instrument can see both ends of this range in a single measurement.20PubMed. The dynamic range problem in the analysis of the plasma proteome Researchers use depletion strategies, essentially stripping away the most abundant proteins before analysis, but these introduce their own complications including lower reproducibility and the risk of inadvertently removing low-abundance proteins that stick to the high-abundance ones.

Beyond the concentration problem, proteins are constantly being made and destroyed. Cells continuously synthesize and degrade proteins to maintain balance and adjust to signals from inside and outside the cell.21PubMed Central. Proteome Turnover in the Spotlight: Approaches, Applications, and Perspectives A snapshot of your proteome at one moment will differ from a snapshot taken an hour later, because some proteins turn over in minutes while others persist for months. The proteome is not a fixed inventory; it is a constantly shifting population.

Emerging single-cell proteomics technologies have begun chipping away at a related challenge: cellular heterogeneity. Even within one tissue, different cell types and cell states express different sets of proteins. Mass spectrometry platforms optimized for single cells have pushed identification past 2,000 proteins per cell, enough to begin distinguishing cell types by their protein profiles.22PubMed Central. Spectral Library-Based Single-Cell Proteomics Resolves Cellular Heterogeneity That is a remarkable technical feat but still covers only a fraction of what each cell likely contains. Efforts like the Human Protein Atlas have mapped protein distribution across tissues, organs, subcellular compartments, blood cells, the brain, and even metabolic pathways, building a spatial picture of where specific proteins show up in the body.23PubMed Central. The Human Protein Atlas-Spatial localization of the human proteome in health and disease

Cancer and the Personalized Protein Problem

If counting proteins in healthy cells is already complicated, cancer makes it worse. Tumors accumulate mutations, and each mutation in a protein-coding region can give rise to mutant peptides called neoantigens, short protein fragments displayed on the surface of cancer cells that look foreign to the immune system. Each missense mutation can potentially generate multiple neopeptides, creating a vast pool of candidate targets, though only a small percentage may trigger a meaningful immune response with any given immune molecule.24PubMed Central. ProGeo-neo: a customized proteogenomic workflow for neoantigen prediction and selection

Modern immunotherapy pipelines now build personalized proteome references for individual patients, annotating every single-nucleotide change and noncanonical transcript in a tumor to predict which peptides might be targetable.25Nature Biotechnology. A comprehensive proteogenomic pipeline for neoantigen discovery to advance personalized cancer immunotherapy The implication is that every person’s tumor carries a partially unique proteome: proteins no other person has, arising from a specific constellation of mutations. This is the extreme end of the “how many proteins” question, where the answer becomes individual-specific.

The Dark Proteome

Even among well-established proteins, a substantial fraction remain structurally mysterious. The so-called “dark proteome” consists of proteins that resist experimental structure determination by current methods and cannot be modeled based on similarity to known structures. A large component of this dark proteome is made up of intrinsically disordered proteins, which do not fold into a fixed three-dimensional shape but instead exist as flexible, fluctuating ensembles.26PubMed. Intrinsically Disordered Proteins: The Dark Horse of the Dark Proteome These disordered proteins are not broken or non-functional. Many play critical roles precisely because their flexibility lets them interact with multiple partners or respond rapidly to signals. But their shapelessness makes them invisible to the crystallography and cryo-electron microscopy techniques that have determined most known protein structures. Recent AI-driven structure prediction tools have made progress on some of these, but the dark proteome remains a significant blind spot in our understanding of the protein universe.

Why Gene Count Is a Poor Proxy for Complexity

One of the more surprising findings from comparative genomics is that humans do not have dramatically more protein-coding genes than organisms that seem far simpler. A roundworm has about 20,000 genes. So does a mustard plant. This apparent paradox dissolves when you look at protein families and functional domains rather than raw gene counts. Humans have roughly 3,300 protein families compared to about 1,800 in the roundworm and about 2,000 in the mustard plant, and a similar pattern holds for the number of distinct protein domains.27PubMed Central. Organismal complexity strongly correlates with the number of protein families and domains The lesson is that complexity comes not from having more genes, but from having a wider toolkit of protein shapes and functions and from the layering of splice variants, modifications, editing events, and regulatory networks on top of them. The human proteome is richer than its gene count suggests, and the raw number of genes is a poor predictor of how complex an organism’s protein world actually is.