How to Calculate and Interpret Fold Change

Fold change is a ratio that describes how much a measured quantity increases or decreases between two conditions, and calculating it is straightforward: divide the value in your experimental group by the value in your control group. A result of 2 means the quantity doubled; a result of 0.5 means it dropped by half. But beneath that simple ratio lies a surprising number of choices about which average to use, whether to work in log space, and how to handle zeros, each of which can meaningfully alter your results and biological conclusions.

The Basic Ratio and Its Asymmetry Problem

At its simplest, fold change equals the measurement in the treated or experimental condition divided by the measurement in the reference or control condition. If a gene’s expression level is 300 units in a treated cell and 100 units in an untreated cell, the fold change is 3. That gene was upregulated threefold. If the numbers are reversed, 100 divided by 300 gives roughly 0.33, meaning the gene was downregulated.

The problem with this linear scale becomes obvious when you try to compare upregulation and downregulation side by side. A threefold increase gives you 3, but a threefold decrease gives you 0.33. Those two changes are the same magnitude in opposite directions, yet on a linear scale one looks enormous and the other looks tiny. This asymmetry makes it hard to visually compare results in tables, plots, or ranked lists, and it’s the main reason most researchers convert fold change into logarithmic form.

Why Log2 Fold Change Is the Standard

Taking the base-2 logarithm of a fold change makes the scale symmetric around zero. A twofold increase becomes a log2 fold change of +1, and a twofold decrease becomes −1. A fourfold increase is +2; a fourfold decrease is −2. The signs tell you the direction, and equal magnitudes in opposite directions sit at equal distances from zero. As one analysis put it, using log2 allows direct interpretation of how many doublings (or halvings) separate the two conditions, with equal fold numbers for up- and downregulation distinguished only by the sign.1PubMed Central. Revisiting Fold-Change Calculation: Preference for Median or Geometric Mean over Arithmetic Mean-Based Methods

This log2 convention is not arbitrary. Biological quantities such as gene expression levels and protein abundances often follow skewed distributions, and logarithmic transformation pulls the long tail in, making the data more symmetric and better suited for standard statistical tests. It also turns multiplicative effects into additive ones, which simplifies downstream calculations considerably. When you see “LFC” or “log2FC” in a paper, this is what it refers to.

Which Average Should You Use?

Fold change compares group-level summaries, so you need to decide how to summarize each group. The instinct is to use the arithmetic mean, the standard average most people learned in school. For normally distributed data, that works fine. But biological measurements are frequently log-normally distributed, meaning most values cluster low with a long tail of high values. In that scenario, the arithmetic mean can be pulled upward by a few extreme readings, distorting the fold change.

A systematic comparison of different calculation approaches found that when data were log-normally distributed, most methods recovered the true ratio between treatment and reference groups as long as the variability in both groups was similar. The exception was the ratio of arithmetic means, which failed even under equal-variance conditions when the two groups had different underlying distributions. All other tested approaches, including ratios of medians and geometric means, proved more robust to unequal variances.1PubMed Central. Revisiting Fold-Change Calculation: Preference for Median or Geometric Mean over Arithmetic Mean-Based Methods

In practical terms, this means that if your data look skewed, you’re better off computing the geometric mean (the nth root of the product of n values) or the median for each group before dividing. The geometric mean is especially natural when you’re already working in log space, since it’s equivalent to exponentiating the arithmetic mean of the log-transformed values. If your data are genuinely normally distributed, the arithmetic mean is fine, but since most high-throughput biological data aren’t, the geometric mean or median is a safer default.

Fold Change in qPCR Experiments

Quantitative PCR, or qPCR, remains one of the most common techniques for measuring gene expression, and it has its own well-established fold change workflow. The dominant approach is the 2^(−ΔΔCt) method, sometimes called the Livak method after the researcher who formalized it. This method provides a convenient way to analyze relative changes in gene expression from real-time quantitative PCR experiments.2PubMed. Analysis of relative gene expression data using real-time quantitative PCR and the 2(-Delta Delta C(T)) Method

The core logic works in two steps. First, you subtract the cycle-threshold value of a reference gene (a “housekeeping” gene that should remain stable across conditions) from the cycle-threshold value of your target gene within the same sample. That gives you ΔCt, which controls for differences in how much total material went into the reaction. Then you subtract the ΔCt of the control condition from the ΔCt of the experimental condition, giving you ΔΔCt. Finally, you raise 2 to the power of −ΔΔCt, and the result is the fold change in expression.

The method’s fundamental logic normalizes all data against the control group to achieve relative quantification of gene expression.3bioRxiv. Comparison of the 2-CT method and the 2-ΔΔCT method for real-time qPCR data analysis One key assumption baked in is that both the target and reference genes amplify with roughly equal efficiency, close to 100%. If one gene doubles every cycle and the other doesn’t, the subtraction step introduces systematic error. Most labs validate this before relying on the method, but it’s worth remembering if your fold changes from qPCR look suspiciously large or small.

Fold Change in RNA-Seq and Count-Based Data

High-throughput sequencing produces count data rather than continuous measurements, and this changes how fold change is estimated. You can’t just divide the raw read count of a gene in one sample by its count in another because sequencing depth, library composition, and technical noise all confound the comparison. Software tools handle this by fitting statistical models that estimate fold change while accounting for these sources of variation.

DESeq2, one of the most widely used tools for RNA-seq differential expression, applies shrinkage estimation to fold changes and dispersion values to improve stability and interpretability.4PubMed Central. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2 In plain terms, shrinkage means that genes with low counts or high variability have their fold change estimates pulled toward zero. Without shrinkage, a gene detected five times in one condition and once in another would show a fivefold change, but that estimate is extremely noisy because the counts are so low. Shrinkage tempers such estimates, reflecting the fact that you’re less certain about the true fold change when the data are sparse.

An alternative approach, GFOLD, was designed specifically for situations where you have only a single biological replicate per condition, which makes standard statistical tests impossible. GFOLD assigns fold change statistics based on a posterior probability distribution rather than a simple ratio, giving more stable rankings of differentially expressed genes even without replicates.5PubMed. GFOLD: a generalized fold change for ranking differentially expressed genes from RNA-seq data If you’re working with limited samples and need to prioritize genes for follow-up, tools like this are worth knowing about.

The Pseudocount Trap

Whenever a gene has zero counts in one condition, taking the log produces negative infinity, and dividing by zero produces an undefined ratio. The standard workaround is adding a small number, called a pseudocount, to all values before computing the fold change. This sounds harmless, but the choice of pseudocount size can substantially alter your results.

Research on pseudocount behavior has shown that they correspond to a specific prior distribution, meaning each choice carries implicit assumptions. When the number of reads for a gene is small, pseudocounts can have substantial influence on the estimated fold change. Using extremely small pseudocounts, like 0.00000001, corresponds to a very wide prior distribution and produces highly exaggerated fold change estimates. But poor choices can also distort fold changes for genes with many reads and affect downstream analyses.6Bioinformatics. Estimating pseudocounts and fold changes for digital expression measurements

The practical lesson is to be deliberate about pseudocounts. Many analysis pipelines add them automatically, and the default value varies between tools. If you’re comparing results across different software packages or publications, differences in pseudocount handling alone can explain discrepancies in reported fold changes for the same genes. When in doubt, check what value was used and consider whether it’s appropriate for the count magnitudes in your data.

Single-Cell RNA-Seq Adds Another Layer of Complexity

Single-cell sequencing introduced a challenge that bulk RNA-seq doesn’t face nearly as severely: dropout. Many genes in any given cell register zero counts not because the gene is truly silent but because the sequencing captured so little material. This makes the pseudocount problem far worse and introduces biases into log fold change estimation that standard tools weren’t designed to handle.

Recent work has shown that widely used single-cell analysis tools like Scanpy and Seurat rely on log-transformed count data with a pseudocount, introducing bias that compromises the reliability of fold change estimates.7bioRxiv. LN’s t-Test: A Principled Approach to t-Testing in Single-Cell RNA Sequencing Newer methods attempt to address this by separately modeling the probability that a gene is expressed at all in a given cell and the mean expression level among cells where it is detected. The result is a less biased fold change estimate with proper confidence intervals. If you’re working with single-cell data, it’s worth checking whether your pipeline’s default fold change calculation has been validated for the dropout rates typical of your protocol.

Ratio Compression in Proteomics

Fold change is used just as widely in proteomics as in genomics, but the mass spectrometry–based methods that quantify proteins introduce their own systematic distortion. Isobaric labeling techniques, which tag proteins from different conditions with chemical labels of the same mass, are known to compress observed fold changes toward the null. In other words, a protein that actually changed threefold might appear to have changed only twofold in the measured data.

A detailed study of this phenomenon found clear bias toward the null for both protein-level and peptide-level analyses. The bias was present across multiple independent experiments, was nearly identical regardless of whether the analysis was done at the protein or peptide level, and was not removed by standard isotope correction procedures. As the true fold change increased, so did the degree of compression.8PubMed Central. Relative Quantification: Characterization of bias, variability and fold changes in mass spectrometry data from iTRAQ labeled peptides This means that in proteomics, the fold changes you see in the data are often underestimates of the real biological changes, and applying the same magnitude thresholds used in RNA-seq would be overly conservative.

Normalization strategies for proteomics data, such as median polishing of log2-transformed reporter ion intensities, help correct for loading and processing differences between channels, but they don’t fully eliminate ratio compression.9EuPA Open Proteomics. Detecting significant changes in protein abundance This is one of the most underappreciated differences between genomic and proteomic fold changes: a 1.5-fold change in a proteomics experiment may be biologically equivalent to a much larger change measured by RNA-seq.

Fold Change Compression in Microarray Preprocessing

Before RNA-seq became dominant, spotted microarrays were the standard platform for gene expression profiling, and they introduced their own version of fold change distortion. Background correction, the step that subtracts non-specific signal from raw intensities, had a direct effect on how accurately fold changes were estimated, especially for low-intensity spots.

A study of microarray preprocessing methods found that most approaches produced fold change compression at low intensities. The methods that minimized compression, specifically the Standard and Edwards background correction approaches combined with a log2 transformation, came with a tradeoff: high variance at low intensities, which in turn produced poor p-value estimates.10PubMed Central. Impact of the spotted microarray preprocessing method on fold-change compression and variance stability Microarray use has declined, but this remains a relevant cautionary tale for anyone working with legacy datasets or reanalyzing older studies. The fold changes reported in a 2005 microarray paper may not be directly comparable to those from a 2023 RNA-seq experiment, partly because of platform-specific compression artifacts.

When Fold Change and Statistical Significance Disagree

A common practice in differential expression analysis is to apply a fold change cutoff alongside a p-value threshold, keeping only genes that pass both. A gene might be statistically significant but change only 1.1-fold, which is often biologically meaningless. Conversely, a gene might show a fivefold change but have so much variance across replicates that the change isn’t statistically reliable. The combination of both filters is what produces the familiar volcano plot, with fold change on the horizontal axis and statistical significance on the vertical axis.

The problem with applying a fold change cutoff as a hard filter is that it ignores the uncertainty around the fold change estimate. A gene estimated at 1.9-fold change is excluded if the cutoff is 2, even though the true change might easily be above 2 given the noise in the data. The TREAT method was developed to address this by formally testing whether a gene’s fold change exceeds a specified threshold, rather than simply filtering on the point estimate. When the magnitude of differential expression is taken into account this way, the approach identifies more biologically relevant genes while improving upon the false discovery rate of existing methods.11PubMed Central. Testing significance relative to a fold-change threshold is a TREAT

If you’re generating a gene list from a differential expression analysis, consider whether you’re using a hard cutoff or a threshold-aware test. The hard cutoff is simpler, but it can both miss real hits and include false ones near the boundary.

When Small Fold Changes Are Biologically Significant

There’s a widespread assumption that only large fold changes matter, that a twofold or greater change is “real” and anything smaller is noise. In many contexts that’s a reasonable heuristic, but it can lead you to dismiss genuinely important biology. Regulatory changes that affect gene expression don’t always need to be dramatic to have large downstream consequences.

A striking example comes from work on three-dimensional genome structure and transcription. Researchers found that subtle changes in the physical contact frequency between regulatory elements, differences so small they were barely visible on standard genomic interaction maps, were sufficient to drive large changes in gene expression. The quantitative disconnect between weak structural boundaries and large transcriptional effects was not an artifact of population averaging; it held up in single-cell analyses.12eLife. How subtle changes in 3D structure can create large changes in transcription The finding challenges the intuition that input fold changes need to be large for output fold changes to be large. In a system with amplification, a small upstream shift can cascade.

This has practical implications for how you set cutoff thresholds. In a signaling pathway with built-in amplification, a 1.3-fold change in an upstream regulator might produce a tenfold change in a downstream target. Filtering out the regulator because its fold change looks unimpressive means missing the cause while catching only the effect.

Comparing Fold Changes Across Platforms

Fold changes measured by different technologies for the same genes often don’t match numerically. RNA-seq, microarray, and qPCR each have platform-specific biases, dynamic ranges, and sensitivities that shift where the measured fold change lands relative to the true biological change. Systematic fold change differences between RNA-seq and microarray data have been documented repeatedly, and ignoring them when combining datasets can inflate false positives or wash out real signals.

A Bayesian approach to integrating microarray and RNA-seq data found that incorporating a normalization step to account for systematic fold change differences between platforms improved detection accuracy and statistical power.13PubMed Central. A Joint Bayesian Model for Integrating Microarray and RNA Sequencing Transcriptomic Data Another study found that concordance between RNA-seq, exon arrays, and qPCR was better for fold change estimates than for raw abundance estimates, suggesting that computing fold change relative to a control is itself a useful normalization step when integrating expression data across platforms.14PubMed Central. Comparative evaluation of isoform-level gene expression estimation algorithms for RNA-seq and exon-array platforms

Rank-based approaches offer another route for cross-platform comparison. Rather than comparing fold change magnitudes directly, which are measured on different scales, some methods convert expression values to ranks within each dataset before looking for agreement. One such tool demonstrated strong accuracy in predicting validated differentially expressed genes across microarray and RNA-seq data.15Nucleic Acids Research. Rank-in: enabling integrative analysis across microarray and RNA-seq for cancer If you’re conducting a meta-analysis across platforms, treating fold change as a relative ranking rather than an absolute magnitude tends to produce more reliable results.

Fold Change in Time-Series Experiments

Most fold change calculations compare two static conditions: treated versus untreated, disease versus healthy, knockout versus wild type. Time-series experiments add a temporal dimension that complicates things. Gene expression doesn’t jump instantaneously from one state to another; it rises and falls in waves, and the maximum fold change for a given gene depends entirely on when you measure it.

One approach to dealing with temporal data is to decompose the expression profile into frequency components, separating slow trends from fast oscillations. A study applying this strategy to a high-resolution Arabidopsis cold-treatment time series found that ranking genes by treatment-frequency fold change produced fewer false positives than the original methodology, and that a specific timepoint (26 hours in their dataset) was the single best statistic for distinguishing known cold-response genes.16PubMed Central. Frequency-based time-series gene expression recomposition using PRIISM The broader point is that in time-series designs, reporting the fold change at a single arbitrary timepoint can be misleading. You either need to justify your chosen timepoint biologically or use an analytical framework designed for temporal data.

For anyone running a time-course experiment, the fold change at the endpoint is not necessarily the most informative number. Genes that spike early and return to baseline may be upstream regulators, while genes that accumulate a fold change slowly may be downstream effectors. Both matter, but they’ll look very different depending on when and how you measure.

Normalization Choices Alter the Fold Changes You Get

Before fold change is even computed, the data typically pass through a normalization step meant to make samples comparable. Different normalization methods rest on different assumptions about the data, and violating those assumptions can introduce systematic bias into every fold change in the dataset.

Most RNA-seq normalization methods assume that the majority of genes are not differentially expressed between conditions. If that assumption holds, the normalization factors correctly adjust for technical differences in sequencing depth and composition. But if a large fraction of genes truly changes, as can happen in aggressive treatments or comparisons between very different cell types, the assumption breaks down, and the normalization can suppress or inflate fold changes globally. The validity of these assumptions has a substantial impact on method performance, and understanding them is necessary for choosing the right method for your data.17PubMed Central. Selecting between-sample RNA-Seq normalization methods from the perspective of their assumptions

In practice, if you’re comparing two conditions where you expect most of the transcriptome to shift, consider adding spike-in controls, synthetic RNA molecules added at known concentrations, that give the normalization algorithm a stable reference. Without them, you’re relying on an assumption that may not hold, and your fold changes may all be biased in the same direction without any obvious sign that something went wrong.