No single housekeeping gene is universally stable across all tissues, treatments, and organisms, which means selecting a reliable reference gene for qPCR normalization requires empirical validation in every new experimental context. Large-scale analyses of gene expression data consistently show that the most popular reference genes, including GAPDH and beta-actin, can vary enough under certain conditions to distort the very results they are supposed to anchor. The good news is that for any given biological context, a subset of genuinely stable genes does exist, and a well-established set of statistical tools can identify them before they compromise your data.
Why the “Classic” Reference Genes Keep Failing
For decades, a handful of genes earned default status in qPCR normalization simply because they were assumed to be constitutively expressed. GAPDH, beta-actin (ACTB), and 18S ribosomal RNA appeared in protocol after protocol without much scrutiny. That assumption has not aged well. A systematic review of endogenous controls in gene expression studies found a striking pattern: the more candidate reference genes a study actually tested, the less likely GAPDH, ACTB, or 18S rRNA were to end up ranked as the most stable choice.1PLOS ONE. With Reference to Reference Genes: A Systematic Review of Endogenous Controls in Gene Expression Studies In other words, these genes looked stable mainly when nobody checked whether anything else was better.
The problem is not that these genes are always bad. It is that they are metabolically active enough to respond to the very conditions researchers study. In a mouse model of intestinal inflammation, for example, GAPDH, ACTB, and beta-2-microglobulin showed the highest variability of all candidates tested, making them poor normalizers in the presence of inflammation.2PLOS ONE. Stability of Reference Genes for Messenger RNA Quantification by Real-Time PCR in Mouse Dextran Sodium Sulfate Experimental Colitis In human brain tumors, ACTB and GAPDH showed different expression levels between low-grade astrocytomas and aggressive glioblastomas.3PubMed Central. ACTB and SDHA Are Suitable Endogenous Reference Genes for Gene Expression Studies in Human Astrocytomas Using Quantitative RT-PCR A gene that shifts between your experimental groups is not normalizing anything; it is introducing a systematic bias that can flip the apparent direction of your target gene’s regulation.
18S ribosomal RNA carries its own issues beyond stability. Because it is a ribosomal RNA rather than a messenger RNA, it does not get captured by oligo-dT priming during reverse transcription, so it requires random priming. Its transcript abundance is also orders of magnitude higher than most target genes, creating a mismatch that can skew quantification. A systematic review in marine bivalve research confirmed that the unvalidated use of 18S rRNA and beta-actin introduced enough artificial variance to change the interpretation of experimental results.4PubMed. Validation of reference genes for RT-qPCR in marine bivalve ecotoxicology: Systematic review and case study using copper treated primary Ruditapes philippinarum hemocytes
How Validation Algorithms Rank Candidate Genes
Rather than guessing which reference gene might be stable, the field has settled on a handful of statistical tools that rank candidates based on expression data from your actual samples. The most widely used are geNorm, NormFinder, BestKeeper, and the comparative delta-Ct method. Each approaches stability from a slightly different angle, and they do not always agree on the final ranking.
geNorm works by comparing every possible pair of candidate genes across all your samples. If two genes track each other closely, rising and falling together, they are considered co-stable. The algorithm identifies the pair with the tightest agreement first, then ranks the remaining genes by how closely they match that top pair. Because it relies on pairwise variation, geNorm always reports two genes tied at the top with the same stability score.5PubMed Central. Optimal use of statistical methods to validate reference gene stability in longitudinal studies This is useful, but it also means geNorm can be fooled by two genes that are co-regulated: they track each other beautifully, but both shift under treatment, giving you a false sense of security.
NormFinder takes a different approach by estimating both within-group and between-group variation for each gene, then combining those into a single stability value. This makes it better at catching genes that look stable overall but actually differ between your treatment and control groups. BestKeeper, meanwhile, uses raw Cq values and their standard deviations to assess stability, then correlates each gene against the geometric mean of all candidates.
Here is where things get tricky for researchers relying on automated platforms. RefFinder is a popular web tool that runs all four algorithms and produces a composite ranking. But a direct comparison found that the BestKeeper output from RefFinder, while producing the same raw data as the standalone BestKeeper software, generated a different final ranking. RefFinder based its BestKeeper ranking on standard deviations of Cq values, while the original BestKeeper software ranks by correlation coefficients against its index. That difference led to substantially different gene rankings.6PLoS ONE. Reference Gene Validation for RT-qPCR, a Note on Different Available Software Packages The lowest correlation between any two methods within RefFinder was between the delta-Ct method and the overall RefFinder composite, suggesting meaningful discrepancies in how different algorithms view the same data.7PLoS ONE. Careful Selection of Reference Genes Is Required for Reliable Performance of RT-qPCR in Human Normal and Cancer Cell Lines
The practical takeaway is not to pick whichever algorithm gives you the answer you like, but to run at least two or three and look for genes that rank well across methods. A gene that tops the list in geNorm but falls to the bottom in NormFinder deserves skepticism. A gene that ranks in the top three across all methods is a much safer bet.
Why You Should Normalize Against Multiple Genes
Even after careful validation, relying on a single reference gene leaves your normalization vulnerable to any remaining instability in that gene’s expression. The foundational paper on this problem demonstrated that the geometric mean of multiple carefully selected housekeeping genes produced a more accurate normalization factor than any single gene alone, a finding validated against large-scale microarray data.8PubMed Central. Accurate normalization of real-time quantitative RT-PCR data by geometric averaging of multiple internal control genes The geometric mean dampens the effect of any one gene wobbling, because the others compensate.
geNorm includes a built-in feature for determining how many reference genes you actually need: it calculates a pairwise variation value (often written as V) for each successive addition of a reference gene to your normalization factor. When adding another gene no longer meaningfully reduces variability, you have enough. For many experiments, two or three validated genes suffice. Using more than necessary adds bench cost without improving precision, while using just one is a gamble. The MIQE guidelines, the community’s agreed-upon reporting standards for qPCR experiments, explicitly emphasize the importance of selecting and validating more than one reference gene for reliable results.9PubMed. Bacterial reference genes for gene expression studies by RT-qPCR: survey and analysis
The MIQE Guidelines and What Reviewers Now Expect
Published in 2009, the MIQE (Minimum Information for Publication of Quantitative Real-Time PCR Experiments) guidelines formalized what many researchers had learned the hard way: without full disclosure of experimental conditions, reagents, primer sequences, and reference gene validation, qPCR results are difficult to reproduce and nearly impossible to evaluate.10PubMed. The MIQE guidelines: minimum information for publication of quantitative real-time PCR experiments The guidelines include a checklist intended to accompany manuscript submissions, covering everything from RNA extraction and reverse transcription conditions to how reference genes were selected and validated.
A companion paper, the MIQE précis, distilled the full guidelines into a practical minimum for day-to-day implementation, including specific guidance on what constitutes an acceptable reference gene validation study versus one that simply picks GAPDH and moves on.11PubMed Central. MIQE précis: Practical implementation of minimum standard guidelines for fluorescence-based quantitative real-time PCR experiments Many journals now require or strongly encourage MIQE compliance, and peer reviewers increasingly flag studies that normalize to an unvalidated single reference gene. If your manuscript does not describe which candidates were screened, which algorithm was used, and how many genes went into the normalization factor, expect questions.
Context Matters More Than Any Gene List
One of the clearest findings across the reference gene literature is that stability is context-dependent, and the contexts that matter are more specific than most researchers initially assume. It is not enough to find genes validated for “human tissue” or “mouse models.” The tissue type, the disease state, the treatment, and even the time course of the experiment can all reshuffle the stability rankings.
Cancer Research
Tumor biology is perhaps the most dramatic example. An analysis using over 10,000 samples from 32 cancer types in The Cancer Genome Atlas confirmed that commonly used reference genes were not consistently expressed across different cancer types.12PubMed Central. Conventionally used reference genes are not outstanding for normalization of gene expression in human cancer research A gene stable in breast tumors might fluctuate wildly in gliomas or colorectal cancers. This is not surprising if you consider that cancer fundamentally reprograms cellular metabolism, altering the expression of genes involved in glycolysis (GAPDH), cytoskeletal dynamics (ACTB), and ribosome biogenesis. But it does mean that any cancer qPCR study needs its own validation, ideally in the specific tumor type and grade being examined.
Stem Cell Differentiation
Stem cell work presents a different challenge: cells changing identity over time. During osteogenic differentiation of human bone marrow stem cells, beta-actin was upregulated with a maximum fold change of about 4.4, likely tied to the morphological changes that come with bone-lineage commitment. GAPDH and RPL13A held up better, with RPL13A showing the tightest stability over the differentiation time course and producing expression patterns of osteogenic markers more consistent with the observed biology.13PubMed Central. Housekeeping gene stability influences the quantification of osteogenic markers during stem cell differentiation to the osteogenic lineage In human pluripotent stem cells undergoing spontaneous differentiation, the picture shifted again: commonly used housekeeping genes varied to a degree that made them inappropriate as references under those conditions.14PubMed. Identification of stable reference genes in differentiating human pluripotent stem cells The lesson is that any experiment involving a change in cell state over time demands longitudinal reference gene validation, not just a snapshot at one time point.
Neurodegenerative Disease
Human brain tissue from patients with Alzheimer’s disease, Parkinson’s disease, and related conditions poses its own normalization challenges. Post-mortem brain samples vary in quality, and the disease process itself alters gene expression in ways that can undermine standard references. Studies evaluating candidates across multiple neurodegenerative conditions identified UBE2D2, CYC1, and RPL13 as the most stable options for brain tissue qPCR work.15PubMed Central. Assessment of brain reference genes for RT-qPCR studies in neurodegenerative diseases These are not household names in qPCR normalization. They emerged only because researchers tested a broad panel in the relevant tissue. Separate validation work in brain tissue from Alzheimer’s, Parkinson’s, and dementia with Lewy bodies patients confirmed that validated reference gene sets need to be disease- and region-specific.16PubMed Central. Identification of valid reference genes for the normalization of RT qPCR gene expression data in human brain tissue
Plant Research Has Its Own Favorites
Animal and plant research share the same normalization principles but often settle on different genes. In plant qPCR, elongation factor 1-alpha (EF1a), beta-tubulin, and protein phosphatase 2A (PP2A) frequently appear among the top-ranked candidates, while ubiquitin-conjugating enzymes and GAPDH show more variable performance depending on the species and stress applied.
In maize, evaluations across abiotic stresses, hormone treatments, and tissue types identified EF1a and beta-tubulin as among the most suitable reference genes, while ubiquitin 9 was flagged as unsuitable.17PLoS ONE. Validation of Potential Reference Genes for qPCR in Maize across Abiotic Stresses, Hormone Treatments, and Tissue Types In creeping bentgrass under salt, drought, cold, and heat stress, the best reference gene pairs shifted with every stress type and even between roots and leaves within the same stress. CACS and PP2A topped the list for heat-stressed roots, but EF1a and UPL7 were the winners in cold-stressed leaves.18PubMed. Selection of reference genes for quantitative real-time PCR normalization in creeping bentgrass involved in four abiotic stresses In fenugreek leaves, different algorithms sometimes converged (EF1a ranked well in both BestKeeper and geNorm) but expression still fluctuated under specific combined stresses like cold plus nanoparticle treatment.19PubMed Central. Characterizing reference genes for high-fidelity gene expression analysis under different abiotic stresses and elicitor treatments in fenugreek leaves
The recurring message from plant work is the same as from animal work, just with a different cast of genes: validate in your species, tissue, developmental stage, and stress condition. A reference gene panel validated for maize drought response tells you nothing about rice under pathogen attack.
When Endogenous References Are Not Enough
Sometimes no endogenous gene stays stable enough to serve as an anchor, especially in experiments where global transcription itself is changing. Viral infections, drug treatments that affect transcription machinery, and massive differentiation events can shift the expression of virtually every gene in the cell. In those cases, exogenous spike-in controls offer an alternative.
The idea is simple: add a known quantity of a synthetic or foreign mRNA to each sample before reverse transcription, then use that spike-in as your reference. Because it is not produced by the organism under study, its measured levels reflect only technical variation in the RNA extraction and reverse transcription steps, not biological changes. Luciferase mRNA has been used this way, and researchers found that it clearly revealed the dynamic expression of endogenous references that would otherwise have been assumed stable, preventing the misinterpretation that comes with a drifting normalizer.20PubMed Central. Exogenous reference gene normalization for real-time reverse transcription-polymerase chain reaction analysis under dynamic endogenous transcription Similarly, an in-vitro-transcribed fragment of the plant gene RuBisCO has been used as a spike-in standard for normalizing qPCR data in bacterial and animal cells, providing a sample-independent alternative.21PubMed. Exogenous reference RNA for normalization of real-time quantitative PCR
Spike-ins have limitations. They do not control for differences in RNA content between cells, because they are added after lysis. They also require careful calibration: add too much and the spike-in dominates the reverse transcription reaction, add too little and the signal is lost in noise. For most standard experimental designs, validated endogenous reference genes remain the better choice. But for experiments where global transcription shifts are the point of the study, spike-ins are worth considering as a complementary strategy.
RNA Quality Quietly Undermines Even Good Reference Genes
Even a perfectly chosen set of reference genes can fail if the RNA going into the reverse transcription step is degraded. An analysis of how RNA integrity affects reference gene performance found that the variation in reference gene expression among the most compromised RNA samples was more than double the variation seen in intact samples. For one representative reference gene (UBC), the noise jumped from about 1.5-fold in high-quality samples to about 3.5-fold in degraded ones. Across 24 combinations of four reference genes and six RNA quality measures, 18 showed a significant difference in reference gene variation between the best- and worst-quality samples.22PubMed Central. Measurable impact of RNA quality on gene expression results from quantitative PCR
Degraded RNA does not affect all transcripts equally. Longer transcripts and those with more fragile secondary structures degrade faster, so the ratio between your reference gene and your target gene shifts in unpredictable ways. Checking RNA integrity before committing to a qPCR run is not just good practice; it directly affects whether your reference genes can do their job. This is another reason the MIQE guidelines place RNA quality assessment early in the checklist.10PubMed. The MIQE guidelines: minimum information for publication of quantitative real-time PCR experiments
A Practical Workflow for Getting This Right
Putting all of this together, a sensible reference gene selection process looks something like this:
- Start broad: Pick at least six to eight candidate reference genes. Include a few traditional ones (GAPDH, ACTB) alongside less common candidates suggested by the literature for your organism and tissue type. Databases like RefGenes, which compile stability data across biological contexts, can help identify candidates with low variance in conditions similar to yours.23BioMed Central / BMC Genomics. RefGenes: identification of reliable and condition specific reference genes for RT-qPCR data normalization
- Run the candidates on your actual samples: Use the same RNA extractions, the same reverse transcription conditions, and the same plates or runs you plan to use for your target genes. Validation on separate pilot samples under different conditions is less informative than it sounds.
- Analyze with multiple algorithms: Run geNorm and NormFinder at a minimum, and consider BestKeeper as a third opinion. Look for genes that rank consistently near the top across methods rather than topping one algorithm and crashing in another.
- Use the geometric mean of two or three top-ranked genes: This is the normalization factor that goes into your delta-Cq or delta-delta-Cq calculations. If your geNorm pairwise variation analysis suggests two genes are sufficient, two is fine.
- Report your validation: Describe which candidates you tested, which algorithms you used, and which genes made the final cut. Journals increasingly expect MIQE-compliant reporting, and a reviewer who sees “normalized to GAPDH” with no validation data will have questions.
The Abundance Mismatch Problem
One issue that gets less attention than stability is the sheer expression level of many traditional reference genes relative to the targets being studied. GAPDH, ACTB, and especially 18S rRNA produce extremely high transcript abundances. When your reference gene has a Cq value of 12 and your target gene comes in at 32, the two are being measured at very different points on the amplification curve, and any small inefficiency differences between the two assays get magnified across that gap. Large-scale transcriptome analyses have confirmed that most commonly used reference genes yield very high transcript abundances compared to the overall transcriptome.23BioMed Central / BMC Genomics. RefGenes: identification of reliable and condition specific reference genes for RT-qPCR data normalization Choosing a reference gene whose expression level is in the same general range as your targets reduces this source of technical error. It is one more reason to move beyond the handful of very-high-abundance defaults and consider moderately expressed candidates like RPL13A, UBE2D2, or CYC1, depending on your system.
This abundance mismatch is also why 18S rRNA, despite being genuinely constitutive in many contexts, can still produce misleading normalization. The gap between its expression and virtually any protein-coding target is enormous, and small pipetting errors or efficiency mismatches create noise that a closer-matched reference would not introduce.