Measuring gene expression means figuring out how actively a cell is reading specific genes and converting them into functional molecules, primarily RNA. The toolkit for doing this ranges from decades-old blotting techniques that look at one gene at a time to modern sequencing platforms that can catalog every active gene in a single cell. Which method you choose depends on whether you need to measure one gene very precisely, survey thousands at once, or preserve the spatial information about where expression is happening inside a tissue. Each approach trades off cost, throughput, sensitivity, and the kind of biological question it can answer.
Measuring One Gene at a Time
The oldest workhorse methods focus on a single gene or a small handful. Northern blotting, developed in the late 1970s, separates RNA molecules by size on a gel, transfers them to a membrane, and then uses a labeled probe to detect a specific transcript. It remains the only method that simultaneously tells you the size of an RNA and how much of it is present across many samples, which is why it still appears in labs studying alternative splicing or transcript processing.1PubMed. Analysis of RNA by Northern Blotting The downsides are obvious: it is slow, it requires a fair amount of starting RNA, and it only looks at whatever gene your probe is designed for.
Quantitative reverse-transcription PCR (RT-qPCR) largely replaced Northern blots for routine quantification. The process starts by converting RNA into a complementary DNA copy, then amplifying a short segment of that copy using gene-specific primers while monitoring the reaction in real time. Because the amplification is exponential, you can calculate how much of the target transcript was present at the start by noting when the signal first rises above background. RT-qPCR can be run in either relative mode, comparing expression between samples, or absolute mode using known standards.2PubMed. Monitoring gene expression: quantitative real-time rt-PCR It is exquisitely sensitive, fast, and relatively cheap per reaction, which is why it remains the go-to method for validating findings from larger-scale experiments and for clinical diagnostics.
One underappreciated wrinkle with RT-qPCR is the choice of reference genes. Every quantification experiment needs an internal control, a gene that is assumed to be expressed at a constant level regardless of the experimental treatment. Selecting a bad reference gene quietly distorts everything measured against it. Researchers use statistical algorithms to rank candidate reference genes by stability, and the best choice often varies depending on the tissue, species, and treatment being studied.3Scientific Reports. Selection and validation of reference genes for quantitative gene expression normalization in Taxus spp. Studies routinely test multiple candidates and find that the classic “housekeeping” genes like actin or GAPDH are not always the most stable option.4PLoS ONE. Using RNA-seq data to select reference genes for normalizing gene expression in apple roots Using the wrong reference gene is one of the most common sources of irreproducible RT-qPCR data, and it is entirely preventable with a bit of upfront validation.
Surveying Thousands of Genes at Once
For questions that require a bird’s-eye view of the entire transcriptome, researchers moved beyond one-gene-at-a-time methods in the 1990s with the invention of DNA microarrays. A microarray works by fixing thousands of known DNA sequences to a surface, then washing fluorescently labeled sample RNA over it. Wherever a sample transcript matches a sequence on the chip, it binds, and the resulting fluorescence indicates how much of that transcript was present.5PubMed Central. Overview of DNA microarrays: types, applications, and their future Microarrays were transformative for cancer research, immunology, and developmental biology throughout the 2000s. Their main limitation is that you can only detect what you already know to look for, since the probes on the chip must be designed in advance.
RNA sequencing, commonly called RNA-seq, solved that limitation. Instead of relying on pre-designed probes, RNA-seq converts the entire pool of RNA in a sample into small complementary DNA fragments, sequences them, and then maps those fragments back to a reference genome. This gives an unbiased, digital readout of every expressed gene, including previously unknown transcripts. RNA-seq has largely supplanted microarrays for most discovery-oriented research because it offers greater dynamic range, can detect novel transcripts, and its cost has plummeted over the past decade. Newer protocols have pushed efficiency even further; some methods now generate transcriptome profiles directly from crude cell lysates in a single tube, eliminating the need for purified RNA and reducing both cost and hands-on time.6Experimental & Molecular Medicine. Cost and time-efficient construction of a 3′-end mRNA library from unpurified bulk RNA in a single tube
Going Deeper with Single-Cell Sequencing
Standard RNA-seq measures the average expression across all the cells in a sample. That average can be misleading when a tissue contains many different cell types, each with its own expression profile. Single-cell RNA-seq (scRNA-seq) solves this by isolating individual cells and sequencing each one separately. Droplet-based microfluidic platforms, which encapsulate single cells in tiny oil droplets along with barcoded beads, have become one of the most widely used isolation strategies because they combine high throughput with relatively low cost per cell.7PubMed Central. Microfluidic design in single-cell sequencing and application to cancer precision medicine A typical experiment can profile tens of thousands of individual cells in a single run.
Not every sample cooperates with standard dissociation, though. Frozen archival tissue, adipose tissue, and certain solid tumors are difficult or impossible to break into intact single cells. For those cases, single-nucleus RNA-seq (snRNA-seq) extracts and sequences nuclei instead. Because nuclear RNA is slightly different from the total cellular pool, the two approaches are complementary rather than interchangeable.8Nature Medicine. A single-cell and single-nucleus RNA-Seq toolbox for fresh and frozen human tumors The availability of snRNA-seq has been especially important for tumor biology, where biobanked frozen specimens vastly outnumber fresh ones.
Seeing Where Genes Are Expressed
Sequencing-based methods, whether bulk or single-cell, require you to physically break apart the tissue first, which destroys the spatial context. If you want to know not just how much of a gene is expressed but where within a tissue section, you need a spatial method. Spatial transcriptomics approaches solve this by capturing RNA directly from intact tissue sections. One foundational strategy places thin tissue slices onto arrays of barcoded capture spots, so that the RNA released from each region of tissue is tagged with a positional barcode. The result is a map of gene expression that can be overlaid on a histological image.9PubMed. Visualization and analysis of gene expression in tissue sections by spatial transcriptomics This type of data has proven valuable in studies of brain organization and tumor heterogeneity, where knowing which genes are active in which region fundamentally changes interpretation.
For even finer resolution, imaging-based methods visualize individual RNA molecules inside cells without any tissue disruption. Single-molecule fluorescence in situ hybridization (smFISH) uses sets of short, fluorescently labeled DNA probes that collectively bind along a target transcript. Each bound transcript appears as a bright diffraction-limited spot under a microscope, and the spots can be counted to give a direct, absolute measurement of copy number with subcellular resolution.10PubMed Central. Optimized protocol for single-molecule RNA FISH to visualize gene expression in S. cerevisiae Multi-color versions of smFISH allow simultaneous detection of several different transcripts in the same cell, and the technique can be combined with immunofluorescence to link RNA levels to protein expression at the single-cell level.11PubMed Central. Single-molecule fluorescence in situ hybridization: quantitative imaging of single RNA molecules The protocol has been adapted to work in intact tissues such as developing fly embryos, where individual mRNA molecules can be localized within complex three-dimensional structures.12PubMed Central. mRNA quantification using single-molecule FISH in Drosophila embryos
Scaling Up In Situ Imaging
Classical smFISH is limited in the number of genes it can measure at once, typically a handful per experiment. A newer generation of multiplexed imaging methods has dramatically expanded that capacity. MERFISH (multiplexed error-robust fluorescence in situ hybridization) assigns each RNA species a unique binary barcode and reads out those barcodes through sequential rounds of hybridization and imaging. A clever error-correction scheme, borrowed from information theory, allows the method to tolerate mistakes in labeling or detection without misidentifying transcripts. Early demonstrations showed imaging of hundreds to a thousand distinct RNA species in individual cells.13PubMed Central. RNA imaging. Spatially resolved, highly multiplexed RNA profiling in single cells Subsequent advances drastically increased throughput, enabling profiling of many more cells per experiment.14PubMed Central. High-throughput single-cell gene-expression profiling with multiplexed error-robust fluorescence in situ hybridization
A related approach called seqFISH uses a different combinatorial labeling strategy, assigning color sequences to each gene across multiple hybridization rounds. Its upgraded version, seqFISH+, can image on the order of 24,000 genes per cell.15PubMed Central. Introduction to bioimaging‐based spatial multi‐omic novel methods These multiplexed imaging methods are converging with sequencing-based spatial transcriptomics to create an increasingly detailed picture of which genes are active in which cells, and exactly where those cells sit within a tissue.
Long-Read Sequencing and Isoform Detection
Standard short-read RNA-seq chops transcripts into small fragments before sequencing, then computationally reassembles the pieces. This works well for counting how much of each gene is expressed, but it struggles to distinguish between different splice variants, or isoforms, of the same gene. Since a single gene can produce multiple isoforms with very different functions, this gap matters. Long-read sequencing platforms from companies like Pacific Biosciences and Oxford Nanopore read full-length transcripts in a single pass, eliminating the need for computational assembly. A large consortium effort recently benchmarked multiple long-read RNA-seq protocols across human, mouse, and manatee datasets, generating over 427 million sequences to evaluate how well different pipelines detect and quantify transcript isoforms.16Nature Methods. Systematic assessment of long-read RNA-seq methods for transcript identification and quantification The technology is still more expensive per base than short-read sequencing, but for questions about splicing, fusion transcripts, or full-length isoform structures, it provides answers that short reads simply cannot.
Measuring What Ribosomes Are Actually Translating
All the methods described so far measure RNA, but RNA levels alone do not tell you how much protein a cell is making from a given transcript. Ribosome profiling bridges that gap by sequencing only the short RNA fragments that are physically shielded by ribosomes during translation. The result is a genome-wide snapshot of which transcripts are being actively translated and where on each transcript ribosomes are sitting, down to single-nucleotide resolution.17PubMed Central. Genome-wide analysis in vivo of translation with nucleotide resolution using ribosome profiling This approach has revealed unexpected complexity in translation, including widespread translation of short open reading frames that were previously dismissed as non-coding, and has become a key tool for understanding translational control.18PubMed Central. Ribosome profiling reveals the what, when, where and how of protein synthesis
A related class of methods focuses on nascent RNA, the transcripts still being synthesized by RNA polymerase. Techniques like GRO-seq and NET-seq capture RNA that is still attached to the transcription machinery, giving a readout of active transcription rates rather than steady-state RNA levels. This distinction matters because a transcript’s abundance at any given moment reflects both how quickly it is made and how quickly it is degraded. Nascent RNA methods disentangle those two processes.19PubMed Central. Comparative analysis of nascent RNA sequencing methods and their applications in studies of cotranscriptional splicing dynamics
The Gap Between RNA and Protein
A persistent complication in gene expression research is that RNA levels do not perfectly predict protein levels. Early large-scale comparisons found only a modest correlation between mRNA and protein abundance, and for years researchers debated how much of the discrepancy was biological and how much was just measurement noise.20PubMed Central. Correlation of mRNA expression and protein abundance affected by multiple sequence features related to translational efficiency in Desulfovibrio vulgaris More refined analyses have pushed the numbers upward. In unperturbed mammalian cells, studies using improved quantification and statistical modeling have estimated that mRNA levels explain roughly 40% to 84% of the variation in protein levels, depending on the analytical approach, with translation rates contributing a relatively small additional fraction.21Cell. How to Measure Gene Expression: An Overview of Methods The current consensus is that at steady state, transcript concentration is the dominant determinant of protein concentration, but the correlation is far from perfect.
In specific organs, the disconnect can be even more pronounced. Kidney studies using the Human Protein Atlas have documented substantial discordance between mRNA and protein expression, with hundreds of proteins detectable by mass spectrometry that were missed by antibody-based methods at the protein level.22PubMed Central. Kidney mRNA-protein expression correlation: what can we learn from the Human Protein Atlas? The practical implication is that gene expression data from RNA-based methods, while enormously informative, should not be treated as a direct proxy for protein function without independent validation.
Reporter Genes and Promoter Activity
Sometimes the question is not “how much RNA is a gene producing right now?” but rather “how active is this gene’s regulatory region?” Reporter gene assays answer that question by fusing a gene’s promoter to a protein that is easy to detect, such as luciferase or a fluorescent protein. When the promoter becomes active, it drives production of the reporter, and the resulting light output or fluorescence serves as a proxy for promoter strength. Dual-luciferase assays, which use a second reporter as an internal control, have become a standard method for characterizing promoter activity both in cultured cells and in whole organisms.23PubMed Central. Application of the dual-luciferase reporter assay to the analysis of promoter activity in Zebrafish embryos
One thing to watch out for with luciferase-based reporters is that the choice of reporter protein matters more than many experimenters realize. Different luciferase variants have different kinetics and sensitivities, and complex biological fluids can interfere with some secreted luciferase systems, complicating interpretation.24Scientific Reports. Reporter gene comparison demonstrates interference of complex body fluids with secreted luciferase activity Reporter assays are powerful for studying transcriptional regulation, but they measure artificial constructs rather than endogenous gene expression, so the results always need to be interpreted with that caveat in mind.
Dealing with Batch Effects
Any large-scale gene expression experiment that processes samples across different days, different reagent lots, or different sequencing runs risks introducing batch effects: systematic technical differences that can masquerade as biological signal. Batch effects are one of the most common sources of false findings in transcriptomics, and correcting for them is a routine but critical step in data analysis. The challenge is that RNA-seq data consists of integer counts that do not follow a simple bell-shaped distribution, so standard statistical corrections can introduce artifacts. ComBat-seq was developed specifically for this problem, using a regression model suited to count data that preserves the integer nature of the measurements.25NAR Genomics and Bioinformatics. ComBat-seq: batch effect adjustment for RNA-seq count data
Even ComBat-seq can lose statistical power in challenging scenarios where batch effects are large. A newer method called ComBat-ref has shown improved performance in simulations, maintaining sensitivity even under conditions where earlier correction methods saw their detection rates drop toward zero.26PubMed Central. Highly effective batch effect correction method for RNA-seq count data For single-cell data, specialized algorithms like scBatch address the additional complexity of correcting batch effects while preserving the ability to identify distinct cell populations.27PubMed Central. scBatch: batch-effect correction of RNA-seq data through sample distance matrix adjustment The field is still actively developing better correction tools, and choosing the wrong one for your data type can do more harm than good.
Gene Expression Signatures in the Clinic
The explosion of transcriptomic data has naturally led to efforts to build clinically useful gene expression signatures, particularly in cancer. The idea is to measure the expression of a defined panel of genes in a patient’s tumor and use the resulting pattern to predict prognosis or guide treatment decisions. Several such signatures have been developed and validated, but the translational path from discovery to clinical use is steep. Despite the large number of proposed prognostic signatures in the literature, only a small fraction have reached the stage of clinical implementation so far.28PubMed Central. Prognostic Cancer Gene Expression Signatures: Current Status and Challenges The reasons include lack of reproducibility across patient populations, overfitting to the discovery cohort, and the difficulty of designing prospective trials to prove a signature actually changes outcomes.
Comparing Expression Across Species
Gene expression measurements are not limited to studying one species at a time. Comparative transcriptomics uses expression data from multiple species to understand how gene regulation has evolved, which developmental programs are conserved, and which have diverged. Modern sequencing has made it feasible to compare entire transcriptomes across organisms, shedding light on the molecular basis of anatomical convergence and innovation.29PubMed Central. What to compare and how: Comparative transcriptomics for Evo-Devo Cross-species comparisons have also clarified the evolution of non-coding RNAs, showing that their sequences and expression patterns change much more rapidly than those of protein-coding genes.30Molecular Biology and Evolution. Comparative Transcriptomics Analyses across Species, Organs, and Developmental Stages Reveal Functionally Constrained lncRNAs
A practical challenge in this work is that many species of interest lack well-annotated reference genomes. The conventional RNA-seq analysis workflow assumes you can map your sequencing reads to a known genome, but for non-model organisms you may need to build a reference transcriptome from scratch before you can even begin counting gene expression. Assembly-free approaches have been developed that bypass this bottleneck by aligning sequencing reads directly to protein databases, enabling differential expression analysis without a reference genome or transcriptome assembly.31PubMed Central. Assembly-free rapid differential gene expression analysis in non-model organisms using DNA-protein alignment These shortcuts make expression profiling accessible for the vast majority of species whose genomes have not been fully sequenced, which is still most of the tree of life.