Underpowered Study: Consequences and Implications for Research

An underpowered study is one that enrolled too few participants (or collected too few observations) to have a reasonable chance of detecting a real effect. The consequences go far beyond simply “missing” a true finding. When underpowered studies do land on a statistically significant result, that result is likely to be inflated, occasionally pointed in the wrong direction, and disproportionately likely to be a false positive. Across large swaths of biomedical and behavioral science, the typical study has far less power than researchers, reviewers, or readers assume, and the downstream damage touches everything from drug development budgets to public trust in scientific findings.

What Power Actually Means and Why It Matters

Statistical power is the probability that a study will correctly detect a real effect when one exists. If a treatment genuinely works, power is your chance of seeing that in the data rather than shrugging and concluding “no significant difference.” By convention, researchers aim for at least 80% power, meaning a one-in-five chance of missing a real effect. In practice, that target is routinely missed. Power depends on three things: the size of the effect you’re looking for, the amount of noise in your measurements, and how many observations you collect.1PubMed. Clinician’s Guide to Understanding Effect Size, Alpha Level, Power, and Sample Size An underpowered study is one where that combination leaves the probability of detecting the effect uncomfortably low.

A common misunderstanding is that low power simply means you’re more likely to get a negative result. That’s true, but it’s the least interesting part of the problem. The more corrosive consequences show up when an underpowered study manages to clear the bar of statistical significance anyway.

The Exaggeration Problem

When a small, noisy study produces a statistically significant result, the estimated effect size is almost guaranteed to be bigger than the true effect. This is sometimes called a Type M (magnitude) error or the exaggeration ratio. The logic is straightforward: in a study with low power, only unusually large random fluctuations will push the result past the significance threshold. The study is, in effect, filtering for flukes. The true effect might be modest, but the only version of it that can survive the statistical test in a tiny sample is an inflated one.2PubMed. Underappreciated problems of low replication in ecological field studies

This is not a niche statistical curiosity. It means that the published literature systematically overstates how large real effects are, because the studies most likely to find “significant” results with small samples are the ones that got lucky with inflated estimates. When other researchers try to replicate the finding with a properly sized sample, they find a smaller effect and may conclude the original was wrong, even though there was a real effect all along. The replication crisis in psychology and biomedicine traces partly to this dynamic.3PubMed. Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors

Getting the Direction Wrong

Even more troubling than exaggeration is the possibility of a Type S (sign) error, where a study not only gets the size wrong but reports the effect as going in the opposite direction from the truth. A drug that slightly helps patients might appear to harm them, or a risk factor that slightly raises disease incidence might look protective. In well-powered studies, sign errors are vanishingly rare. But as power drops, their probability climbs. One analysis showed that as power decreased, the rate of significant results pointing in the wrong direction increased, and the magnitude of those wrong-direction estimates grew larger as well.4PubMed Central. Statistically significant results from low-power analyses: A comedy of errors

For a field trying to build cumulative knowledge, this is a nightmare. A handful of underpowered studies producing sign errors can send other researchers chasing a relationship that not only doesn’t exist in the reported magnitude but doesn’t even go in the reported direction. Gelman and Carlin’s framework for assessing Type S and Type M errors before running a study has gained traction as a practical planning tool, but it remains far from standard practice.3PubMed. Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors

False Discoveries Pile Up

Low power also inflates the false discovery rate, meaning that a larger share of “significant” findings in a body of literature are statistical artifacts rather than real effects. When you combine weak power with the conventional 5% significance threshold, the math gets ugly. A simulation-based analysis found that, under conditions reflecting current practice in many subfields of biomedical science, at least 25% of reported discoveries may be false. The authors argued that the only viable corrective is to enforce a minimum of 80% power for published studies.5PubMed. How often should we expect to be wrong? Statistical power, P values, and the expected prevalence of false discoveries

In psychology specifically, a study tracking trends from 1975 to 2017 found that the combination of low power and selective reporting practices meant a considerable share of significant results were likely just statistical artifacts.6PLOS ONE. Are most published research findings false? Trends in statistical power, publication selection bias, and the false discovery rate in psychology (1975–2017) The problem isn’t dishonesty. It’s that the system is set up so that noise, filtered through underpowered designs and publication incentives, mimics signal.

How Widespread Is the Problem

If you’re hoping that low power is confined to a few careless labs, the data are discouraging. A large empirical assessment of studies in cognitive neuroscience and psychology found that median power to detect small effects was just 12%, and even for medium-sized effects it was only 44%. Power to detect large effects, the easiest targets, reached 73%, still short of the 80% benchmark. These numbers showed no improvement over the preceding half century.7PubMed Central. Empirical assessment of published effect sizes and power in the recent cognitive neuroscience and psychology literature

Neuroscience fares even worse. An influential review estimated median power in neuroscience studies at somewhere between 8% and 31%, depending on the subfield and the effect being tested.8PubMed Central. Power failure: why small sample size undermines the reliability of neuroscience A separate review of three human-research domains in biomedical science found that roughly half of studies had power in the 0–20% range. Median power ranged from less than 9% in Alzheimer’s disease research to about 30% in studies of major depressive disorder, with somatic diseases, psychiatric disorders, and neurological diseases all hovering around 17–20%.9PubMed Central. Low statistical power in biomedical science: a review of three human research domains

These aren’t fringe subfields. Alzheimer’s research, depression treatment, oncology, cardiology: some of the areas where patients most need reliable answers are the same areas drowning in underpowered studies. And because many of the true effects in these fields are small, low power is especially damaging. The smaller the real effect, the more participants you need, and the bigger the exaggeration when an underpowered study stumbles onto significance.

The Ethics of Enrolling People in Underpowered Trials

Clinical trials aren’t just data-collection exercises. They involve real people accepting real risks, from side effects to the possibility of receiving a placebo instead of active treatment. When a trial is too small to have a reasonable chance of detecting whether the treatment works, those participants are exposed to risk with little prospect of the study producing useful knowledge. An analysis in JAMA argued that this makes most underpowered clinical trials unethical and concluded that they’re justifiable in only two narrow situations: small trials for rare diseases where investigators have explicit, pre-specified plans to pool their results in a prospective meta-analysis, and early-phase trials designed for purposes other than randomized treatment comparisons. In both cases, participants must be told that their involvement may only indirectly benefit future patients.10PubMed. The continuing unethical conduct of underpowered clinical trials

That second condition matters. Not every underpowered trial is a pointless waste of participants’ goodwill. Dose-finding studies, safety signal studies, and proof-of-concept work serve legitimate purposes without needing to be powered for definitive efficacy testing. The ethical problem arises when a trial claims to test efficacy but was never large enough to do so, a situation that remains common.11Ethics, Medicine and Public Health. Ethical concerns of including too few or too many participants in clinical studies

Underpowered Studies Contaminate Meta-Analyses

A popular defense of small studies is that they can be “rescued” by future meta-analyses, which combine results across many studies to arrive at a more reliable estimate. This defense has some truth but carries a serious catch: if the individual studies feeding into a meta-analysis are systematically underpowered, the meta-analysis inherits their biases. Underpowered studies are more likely to go unpublished if they find null results (publication bias), and the ones that do get published are the ones with inflated effect sizes. The meta-analysis then aggregates a skewed sample of the evidence.12PubMed Central. The Perils of Misinterpreting and Misusing “Publication Bias” in Meta-analyses: An Education Review on Funnel Plot-Based Methods

Making things worse, the statistical tests commonly used to detect publication bias within a meta-analysis are themselves underpowered when the number of included studies is small. An analysis of Cochrane review meta-analyses confirmed that standard tests underestimate the presence of publication bias, particularly when few studies are available.13PubMed. P value-driven methods were underpowered to detect publication bias: analysis of Cochrane review meta-analyses So the very tool designed to correct for low-quality small studies is itself compromised by the same problem. The chain of error propagates upward through the evidence hierarchy.

Why Researchers Keep Running Small Studies

If the problems are this well-documented, why hasn’t behavior changed? The answer is largely structural. An optimality model of researcher incentives found that scientists acting rationally to maximize their publication records and career success should spend most of their effort seeking novel results and conducting studies with only 10–40% statistical power.14PubMed Central. Current Incentives for Scientists Lead to Underpowered Studies with Erroneous Conclusions That’s not a failure of individual integrity. It’s the predictable output of a system that rewards publication volume and novel findings over replication and rigor.

The logic, from any individual researcher’s perspective, is rational even if collectively destructive. Running three small studies gives you three chances at a significant, publishable result. Running one large, properly powered study ties up the same resources in a single bet. The underpowered approach benefits the individual researcher’s career while harming the field’s collective ability to accumulate reliable knowledge.15Social Psychological and Personality Science. A Powerful Nudge? Presenting Calculable Consequences of Underpowered Research Shifts Incentives Toward Adequately Powered Designs Funding structures reinforce this: grants often provide enough money for one moderately-sized study, and researchers stretch those budgets across multiple smaller projects.

The Financial Toll

Underpowered studies don’t just waste participants’ time and goodwill. They waste money on a staggering scale. An analysis of irreproducibility in preclinical life science research estimated that more than half of preclinical studies are not reproducible, and that this irreproducibility costs roughly $28 billion per year in the United States alone.16PubMed Central. The Economics of Reproducibility in Preclinical Research Low statistical power is one of the major drivers of irreproducibility, alongside flawed reagents, poor study design, and selective reporting. When an underpowered preclinical study produces an inflated effect size that can’t be replicated, downstream clinical development programs burn through years and millions of dollars before discovering the original finding doesn’t hold up.

The Post Hoc Power Trap

One common response to criticism of an underpowered study is to compute power after the fact, using the observed effect size and sample size from the completed study. This “post hoc” or “retrospective” power analysis is widely considered misleading. Simulation work has shown that post hoc power estimates are highly variable in the range researchers care about most and can differ dramatically from true power. The reason is almost circular: in a study that found a non-significant result, observed post hoc power will always be low (because the observed effect was small or noisy), telling you nothing you didn’t already know from the non-significant p-value. And in a study that found a significant result, post hoc power will always look adequate, even if the study was genuinely underpowered and just got lucky.17PubMed Central. Post hoc power analysis: is it an informative and meaningful analysis?

The upshot is that power calculations are useful before data collection but essentially useless afterward. If you see a published paper defending itself with a post hoc power analysis, treat that defense with skepticism. The appropriate time to think about power is during study design, not in the discussion section.

Registered Reports and Other Structural Fixes

If the problem is structural, the solutions have to be structural too. One of the most promising innovations is the registered report, in which a study’s design, hypotheses, and analysis plan are peer-reviewed before data collection begins. The journal commits to publishing the results regardless of outcome, which strips away the incentive to run small studies hoping for lucky significant results. Registered reports help prevent low statistical power, selective reporting, and publication bias in a single mechanism.18PLOS ONE. Recommendations in pre-registrations and internal review board proposals promote formal power analyses but do not increase sample size

However, the evidence on their impact is nuanced. Making power analyses a required part of pre-registration or ethics board proposals does increase the number of researchers who conduct formal power analyses. But that same research found it didn’t necessarily increase sample sizes, suggesting that researchers sometimes go through the motions of a power analysis without letting it change their study design. A review focused on preclinical neuroscience argued that registered reports are the most immediately impactful strategy for improving replicability, but acknowledged that adoption remains limited and that changing the broader incentive landscape is ultimately necessary.19eNeuro. Questionable Research Practices, Low Statistical Power, and Other Obstacles to Replicability: Why Preclinical Neuroscience Research Would Benefit from Registered Reports

Big Team Science and Resource Pooling

Another structural fix attacks the problem from the supply side: if individual labs can’t afford large samples, why not distribute data collection across many labs? “Big team science” collaborations pool resources from multiple research groups, generating sample sizes that no single lab could manage alone. This approach is especially valuable when studying small effects, testing complex interactions, or running analyses that require corrections for multiple comparisons.20PubMed Central. How to build up big team science: a practical guide for large-scale collaborations

Large-scale replication projects in psychology have used this model to test high-profile findings with samples vastly larger than the originals, and the results have been humbling: effect sizes in the replications are consistently smaller than in the original publications, exactly the pattern you’d predict if the originals were underpowered and publication-biased. Big team science isn’t a panacea; it’s expensive, logistically challenging, and doesn’t work for every type of research question. But for confirmatory studies of effects that matter, it directly addresses the sample-size bottleneck.

Bayesian Approaches as a Complement

Some researchers argue that the entire framework of null-hypothesis significance testing, where power is defined, is part of the problem. Bayesian statistical methods offer an alternative in which evidence is accumulated incrementally, prior knowledge from earlier studies can be formally incorporated, and researchers can directly test whether an effect is zero rather than merely failing to reject a null hypothesis. Bayesian estimation allows results from all previous research to be combined with study estimates in a principled way, sometimes yielding support for hypotheses that frequentist methods miss.21Journal of Management. Bayesian Estimation and Inference

Bayesian methods don’t magically solve the small-sample problem. A tiny, noisy dataset will produce wide posterior distributions regardless of the statistical framework. But Bayesian approaches do change the conversation: instead of asking “is the p-value below 0.05?” you’re asking “how much should we update our beliefs given this evidence?” That shift can reduce the all-or-nothing thinking that leads researchers to chase significance in underpowered designs. Where prior evidence exists, Bayesian frameworks let a small study contribute to cumulative knowledge without pretending to be definitive on its own.

Misunderstandings That Keep the Problem Alive

Misinterpretation of statistical concepts is not a minor pedagogical gap; it’s an active contributor to the underpowered-study problem. Researchers, reviewers, and even textbooks routinely promote shortcut definitions of p-values, confidence intervals, and power that are simply wrong. One comprehensive guide documented dozens of persistent misinterpretations, noting that correct use of these concepts demands an attention to detail that seems to tax the patience of working scientists, leading to an epidemic of incorrect shortcut definitions that dominate the literature.22PubMed Central. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations

For example, many researchers believe that a non-significant result from an underpowered study means “the treatment doesn’t work.” It doesn’t. It means the study couldn’t tell. Others believe that a significant result from an underpowered study is especially impressive because it overcame long odds. In reality, as covered earlier, it’s more likely to be exaggerated or false. Still others compute post hoc power and believe they’ve assessed study adequacy, when as noted they’ve really just performed an algebraic transformation of their p-value. Each of these misunderstandings leads to real decisions: wrong treatments pursued, right treatments abandoned, resources allocated based on noise dressed up as signal.

What Changes When You Read Studies With Power in Mind

If you regularly read research, whether for clinical decision-making, policy work, or personal curiosity, keeping power in mind changes what you trust. A striking result from a small study should raise your eyebrows, not lower them. The splashier the finding and the smaller the sample, the more likely the effect is inflated. Conversely, a null result from a small study tells you very little; the absence of evidence is not evidence of absence when the study was never equipped to find anything.

When reading a meta-analysis, check whether the included studies are individually underpowered. If most have sample sizes in the dozens, the pooled estimate may inherit inflated effect sizes from publication bias. Look for whether the meta-analysis tested for publication bias and whether that test itself had enough studies to be reliable. And when someone defends a questionable finding by pointing to a post hoc power analysis, you now know why that defense doesn’t hold water. The tools for evaluating evidence quality are not complicated once you know where the weaknesses lie. The hard part is resisting the pull of a clean, significant p-value from a study that was never designed to produce one honestly.

Leave a Reply

Your email address will not be published. Required fields are marked *