Sample size determines whether a study can reliably detect what it set out to find. Too few participants and a real effect can slip through unnoticed, producing a false negative. But the problem runs deeper than missed findings: studies with inadequate sample sizes are also more likely to produce false positives, inflate the apparent size of effects, and fail when other researchers try to replicate them. Getting the number of participants right is one of the most consequential decisions in any study’s design, and the consequences of getting it wrong ripple through entire fields of science.
How Small Samples Undermine Reliability
When researchers design a study, they want to know whether something they’re testing actually works or actually exists. A drug might lower blood pressure, a teaching method might improve test scores, or a brain region might be linked to a behavior. The study’s ability to detect a real effect like this is called its statistical power. When the sample is too small, power drops, and the study becomes less likely to pick up on something that’s genuinely there. That much is intuitive. What’s less obvious is the second problem: when an underpowered study does manage to produce a statistically significant result, that result is less likely to reflect something real.1PubMed Central. Power failure: why small sample size undermines the reliability of neuroscience
Think of it this way. Imagine you flip a coin ten times and get eight heads. You might start wondering if the coin is rigged. But if you flip it a thousand times and get 520 heads, the slight bias is both more detectable and more trustworthy because the larger sample smooths out the randomness that dominates small samples. A small study is like those ten flips: any pattern you see could easily be noise, and the patterns that do cross the threshold of “statistically significant” tend to be exaggerated versions of what’s actually happening.
One analysis of biomedical research estimated that many subfields can expect their discoveries to be wrong at least a quarter of the time, largely because of insufficient power. The proposed fix was straightforward: require that studies be designed with at least 80% power, meaning they have an 80% chance of detecting a true effect of the expected size.2PubMed. How often should we expect to be wrong? Statistical power, P values, and the expected prevalence of false discoveries That threshold has become a widely cited benchmark, though plenty of published studies still fall short of it.
Precision Gets Tighter as Numbers Grow
Sample size also controls how precisely a study can estimate a quantity. When researchers report a finding, they usually attach a confidence interval around it: a range that plausibly contains the true value. A small study produces a wide interval because there’s more uncertainty in the estimate. A larger study narrows that range, giving a more useful answer. The width of a confidence interval depends on the sample size and the variability in whatever is being measured.3PubMed Central. Using the confidence interval confidently
This matters practically. If a clinical trial reports that a new treatment reduces symptoms by somewhere between 2% and 40%, that range is so wide it’s almost useless for deciding whether the treatment is worth using. A larger trial might narrow that to 15% to 25%, which is much more actionable. The point estimate might be similar in both cases, but the precision of the larger study makes the finding far more informative. Researchers and clinicians care not just about whether an effect exists, but about how big it is and how confident they can be in that estimate.
The Reproducibility Connection
One of the most discussed problems in modern science is the failure of studies to replicate. A high-profile finding gets published, another lab tries to repeat the experiment, and the result vanishes. Small sample sizes have been identified as a major driver of this problem. In neuroimaging, for instance, studies linking brain features to behavioral traits have historically used samples of around 25 people. Research published in Nature found that these associations were far smaller than previously thought, and that reproducible results only emerged when sample sizes grew into the thousands. At typical sample sizes, both inflated effect sizes and outright replication failures were common.4PubMed Central. Reproducible brain-wide association studies require thousands of individuals
That finding forced a reckoning in the neuroimaging community. Many influential brain-behavior studies had been built on samples that simply could not support the claims being made. The effects were real, but they were small, and small effects need large samples to be reliably detected. When a field’s standard sample size is orders of magnitude below what the actual effect sizes demand, replication failures aren’t surprising. They’re inevitable.
There’s an interesting wrinkle here, though. A large-scale analysis pooling over 300 original-replication study pairs from major replication projects found that the original study’s sample size did not, by itself, predict whether it would replicate. What did predict replicability was the size of the effect being studied. Studies investigating larger effects replicated more often, regardless of how many participants were involved. The reason the two are connected is that studies with bigger samples tended to investigate smaller effects, which are inherently harder to replicate.5PubMed Central. Challenging the N-Heuristic: Effect size, not sample size, predicts the replicability of psychological science In other words, sample size matters because it needs to be calibrated to the effect you’re trying to detect. A study of 50 people looking at a huge effect can be perfectly reliable. A study of 50 people looking at a tiny effect cannot.
When Too Many Participants Is Also a Problem
Most conversations about sample size focus on the dangers of too few participants, but too many can create its own distortions. With a sufficiently large sample, even trivially small effects become statistically significant. A drug might lower blood pressure by a fraction of a point that has no meaningful clinical impact, but with enough data points the difference passes the significance threshold and gets reported as a “finding.” The study is not wrong in a statistical sense, but it is misleading in a practical one. This disconnect between statistical significance and real-world importance becomes more pronounced as datasets grow larger.
This is increasingly relevant as big datasets become the norm in fields like genomics, electronic health records research, and social science. When you’re analyzing millions of records, almost everything is statistically significant. The discipline shifts from asking “is there an effect?” to asking “is the effect big enough to matter?” Researchers working with very large datasets need to define in advance what size of effect would be clinically or practically meaningful, and focus on that rather than chasing statistical significance alone.
The Ethics of Getting It Right
Sample size is not just a statistical concern; it’s an ethical one. Clinical trials expose participants to experimental treatments with uncertain benefits and real risks. If a trial enrolls too few people and produces inconclusive results, those participants bore the risks and burdens of participation for nothing. The data doesn’t contribute useful knowledge, and the question the trial set out to answer remains unanswered. Guidelines for clinical research recognize this: a sample too small to yield valuable results means participant exposure was not justified.6JAMA. Ethical Considerations for the Inclusion of Patient-Reported Outcomes in Clinical Research
The reverse also applies. Enrolling more participants than necessary exposes extra people to potential harms and wastes resources that could support other research. In some situations, ethical review boards specifically require unequal group sizes. For placebo-controlled trials involving seriously ill patients, for example, assigning equal numbers to the placebo arm and the treatment arm can be ethically problematic, and researchers may be required to put more participants in the treatment group.7PubMed Central. Sample Size Estimation in Clinical Trial – Section: Importance of Sample Size in CT The goal is to enroll exactly the number needed to answer the question with adequate confidence and not a person more.
What Happens When Large Samples Are Impossible
Not every research question comes with the luxury of thousands of willing participants. Rare diseases, by definition, affect small populations. A condition that shows up in one in a hundred thousand people means there simply aren’t enough patients to fill a conventionally powered trial. Standard methods for detecting minimum clinically important differences often require sample sizes that are completely unfeasible for these populations.8PubMed. Design and analysis features used in small population and rare disease trials: A targeted review
Researchers working in these areas have developed workarounds. Adaptive trial designs allow the protocol to change as data accumulates, reallocating participants to more promising treatment arms. Bayesian statistical methods incorporate prior knowledge rather than starting from scratch, which can squeeze more information from fewer data points. Crossover designs, where each participant serves as their own control by receiving both the treatment and the placebo at different times, effectively double the informational value of each person enrolled. N-of-1 trials, which systematically test treatments in a single patient, represent the extreme end of this spectrum. None of these fully replace the assurance that comes from a large, well-powered trial, but they represent the best available approach when the population simply doesn’t exist to support one.
Measurement Quality Changes the Equation
One underappreciated factor in sample size planning is the quality of the measurements being taken. If the tools used to collect data are imprecise or inconsistent, the noise in the data increases, and a larger sample is needed to see through that noise. Modeling work has shown that improving measurement reliability from 0.7 to 0.9 (on a scale where 1.0 is perfect) can reduce the required sample size by about 22%.9PubMed. Penny-wise and pound-foolish: the impact of measurement error on sample size requirements in clinical trials That’s a substantial saving in time, money, and participant burden, achieved not by enrolling more people but by measuring more carefully.
The flip side is equally important. If researchers ignore measurement error when planning their study, they’ll underestimate how many participants they need. This can lead to studies that look adequately powered on paper but are effectively underpowered because the data is noisier than assumed.10PubMed. The impact of ignoring measurement error when estimating sample size for epidemiologic studies Investing in better training for data collectors, using standardized instruments, or averaging multiple measurements of the same thing can all improve reliability and reduce the number of participants needed. Researchers sometimes fixate on enrolling more people when cleaning up their measurement process would accomplish more.
Planning the Right Number Before the Study Starts
Responsible study design involves calculating the required sample size before enrollment begins. This calculation depends on several inputs: the expected size of the effect you’re trying to detect, the variability in your outcome measure, the significance threshold you’ll use, and how much power you want. The effect size is often the hardest part, because it requires estimating what you’re going to find before you find it. Researchers typically draw on pilot studies, previous literature, or expert judgment to make this estimate, and the choice of effect size essentially determines the required sample size once the other parameters are set.11PubMed. The importance of a priori sample size estimation in strength and conditioning research
Getting this wrong has cascading consequences. If you overestimate the likely effect, you’ll plan for too few participants and end up underpowered. If you underestimate, you’ll enroll more people than necessary. The tools for doing these calculations are widely available and relatively straightforward to use.12PubMed Central. Sample size determination and power analysis using the G*Power software The challenge is usually not the arithmetic but the judgment calls that feed into it.
Budget constraints add another layer. In cluster-randomized trials, where groups of people rather than individuals are assigned to conditions, the optimal sample size depends not just on statistical parameters but also on the costs of sampling and measuring each cluster and each person within a cluster. Researchers must sometimes trade off between statistical ideals and practical realities, finding the design that maximizes power for a fixed budget rather than the budget needed for ideal power.13PubMed. Sample size calculation in cost-effectiveness cluster randomized trials: optimal and maximin approaches
Subgroup Analyses Demand Even More Participants
A study might be well-powered to detect an overall treatment effect but still lack the numbers to answer important secondary questions. When researchers want to know whether a treatment works differently in men versus women, in older versus younger patients, or across racial and ethnic groups, each of those subgroup comparisons requires its own statistical power. A study powered to detect an effect in the full sample may be hopelessly underpowered to detect effects within subgroups, especially if the subgroups are unequal in size. Methods for calculating sample sizes to test subgroup-specific effects have not received as much attention as methods for testing overall effects, despite growing interest in health equity research.14PubMed Central. Sample Size Requirements to Test Subgroup-Specific Treatment Effects in Cluster-Randomized Trials
This matters because subgroup differences are often the most clinically important findings. A treatment that works on average but harms a particular demographic is not a success story. Yet the sample sizes needed to reliably detect subgroup-specific effects can be several times larger than what’s needed for the overall analysis. Funders and review boards are increasingly recognizing that diversity in trial enrollment is not just a fairness issue but a statistical requirement for answering the questions that matter most.
How Meta-Analysis Compensates for Individual Study Limitations
When individual studies are small, researchers can pool their results through meta-analysis. By combining data from multiple studies addressing the same question, a meta-analysis can achieve the statistical power that no single small study could provide on its own. It can also identify patterns across studies, such as whether a treatment works better in certain populations or settings.15PubMed. Pooling research results: benefits and limitations of meta-analysis
But meta-analysis has its own vulnerabilities. The most insidious is publication bias: studies that find dramatic, statistically significant results are more likely to get published than studies that find nothing. If the published literature is skewed toward positive findings, a meta-analysis that pools those studies will inherit the skew and overestimate the true effect. Methods exist for detecting and quantifying publication bias, such as examining whether the collected studies’ distribution is asymmetric in ways that suggest missing null results.16PubMed Central. Quantifying publication bias in meta-analysis But correcting for it after the fact is imperfect. The best defense against publication bias is conducting adequately powered individual studies and reporting their results regardless of the outcome.
Sample Size in the Age of Machine Learning
Machine learning introduces a different dimension to the sample size question. Traditional statistical tests ask “is there an effect?” while machine learning algorithms ask “can I predict the outcome for a new case I’ve never seen?” The sample sizes needed for building reliable predictive models depend on the complexity of the data, the number of variables, and how strong the underlying signal is. In genomics, for example, classification tasks using RNA sequencing data required median sample sizes ranging from roughly 190 to 480 individuals depending on the algorithm used, with stronger biological signals and more balanced classes reducing the requirement.17PubMed Central. Sample size requirements for machine learning classification of binary outcomes in bulk RNA-Seq data
A useful rule of thumb for high-dimensional biomarker panels is that the required number of cases per variable depends on both the predictive strength of the panel and how concentrated the signal is among the variables. When most biomarkers contribute something to prediction, you need more data per variable than when only a few biomarkers carry the signal. To learn a classifier that captures 80% of the available predictive information, the required cases per variable can range from as few as 0.1 for strong, sparse signals to about 9 for weaker, diffuse ones.18PubMed. Sample size requirements for learning to classify with high-dimensional biomarker panels The numbers vary enormously with context, which means off-the-shelf rules like “you need ten observations per variable” are crude at best.
Why Sample Sizes Haven’t Grown as Fast as the Advice
Given how well understood the problems of small samples are, you might expect that studies have gotten bigger over time. The trend is less encouraging than you’d hope. A comparison of sample sizes across leading psychology journals in 1955, 1977, 1995, and 2006 found that recommendations for increasing sample sizes had not been meaningfully integrated into core psychological research.19PubMed. Sample size in psychological research over the past 30 years The picture varies somewhat by subfield, but the overall pattern is one of slow progress despite decades of warnings from methodologists.
Several forces conspire to keep sample sizes low. Running participants is expensive and time-consuming. Grant funding often doesn’t cover the cost of adequately powered studies, especially for effects that turn out to be smaller than initially expected. Researchers face pressure to publish frequently, which favors running many small studies over fewer large ones. And the traditional publishing ecosystem has rewarded novel, significant findings over rigorous null results, which creates an incentive structure that tolerates underpowered work as long as it produces an eye-catching result. Reform efforts, including preregistration requirements, registered reports that commit journals to publishing results regardless of outcome, and funder mandates for power calculations, are slowly changing the landscape. But the gap between what methodologists recommend and what most researchers practice remains wide.