How to Establish Causation in Scientific Research

Establishing causation in scientific research requires ruling out alternative explanations for why two things appear connected. A correlation between ice cream sales and drowning rates does not mean ice cream causes drowning; both rise together because of summer heat. The core challenge in any causal claim is demonstrating that X actually produces Y, not that they merely travel together. Researchers have developed a toolkit of study designs, logical frameworks, and statistical methods to do exactly that, and the right tool depends heavily on the question being asked.

Why Correlation Falls Short

The gap between correlation and causation is not just a slogan. Two variables can move in lockstep for reasons that have nothing to do with one affecting the other. The most common culprit is a third, unmeasured factor driving both. In research, this is called confounding. When two variables correlate not because one directly causes the other but because a third variable brings about both, the observed relationship is spurious.1Encyclopedia of Measurement and Statistics. Spurious Correlation A famous historical example: countries with higher chocolate consumption also win more Nobel Prizes. The real explanation is national wealth, which funds both confectionery markets and research universities.

Confounding is not the only trap. Sometimes the direction of the relationship is backward. Researchers studying whether depression causes unemployment might actually be observing the reverse: job loss triggering depression. Disentangling which way the arrow points, known as reverse causality, is a persistent challenge, particularly when both variables feed on each other over time.2Journal of Developmental and Life-Course Criminology. Reciprocal Relationships, Reverse Causality, and Temporal Ordering: Testing Theories with Cross-lagged Panel Models These problems are why establishing causation demands specific study designs rather than simply collecting data and looking for patterns.

The Randomized Controlled Trial

If you can randomly assign people to receive a treatment or a placebo, you have the strongest tool available for causal inference. The logic is simple: when participants are allocated randomly, every trait that might influence the outcome, whether the researchers know about it or not, gets distributed roughly evenly between groups. Any difference that emerges can be attributed to the treatment itself rather than some lurking variable. This is why randomized, double-blind, placebo-controlled trials are widely considered the gold standard for demonstrating that an intervention causes an outcome.3PubMed Central. Randomized double blind placebo control studies, the “Gold Standard” in intervention based studies

Blinding adds another layer. When neither the participant nor the researcher knows who got the real treatment, the results cannot be skewed by expectations. A patient who knows they received the active drug might report feeling better just because of that knowledge. A clinician who knows might unconsciously evaluate symptoms more favorably. Double blinding eliminates both of these biases.

But randomized trials have real limits. You cannot randomly assign people to smoke for twenty years to see if it causes lung cancer. You cannot randomly assign poverty. Many of the most important questions in public health, social science, and economics involve exposures that are either unethical or impossible to assign by coin flip. Randomized trials also tend to study narrow populations under controlled conditions, which can limit how well their findings apply to the messier real world.4PubMed Central. Causal Inference Methods for Combining Randomized Trials and Observational Studies: A Review A drug tested in carefully screened adults aged 40 to 65 with no other health conditions may behave differently in an 80-year-old with three chronic diseases.

The Bradford Hill Framework for Observational Evidence

When experiments are not feasible, researchers working with observational data need a systematic way to weigh whether an association is likely causal. The most influential checklist for this purpose comes from a 1965 lecture by the epidemiologist Austin Bradford Hill. His nine viewpoints remain the most frequently cited framework for causal inference in epidemiology.5PubMed Central. Applying the Bradford Hill criteria in the 21st century: how data integration has changed causal inference in molecular epidemiology They are not a checklist where every box must be ticked, but rather a set of considerations that strengthen or weaken the case for causation.

The viewpoints include ideas that feel intuitive once named:

  • Strength: A large effect is harder to explain away by confounding alone. Smokers having fifteen times the lung cancer risk of nonsmokers is harder to dismiss than a 10% difference.
  • Consistency: The same association shows up across different populations, settings, and study designs.
  • Temporality: The supposed cause must precede the effect. This is the one viewpoint Hill considered non-negotiable.
  • Dose-response: More exposure leads to more effect. Heavier smoking predicts higher cancer risk.
  • Plausibility: A plausible biological or mechanical explanation exists for how X could cause Y.
  • Coherence: The causal interpretation does not conflict with what is known about the disease or phenomenon.
  • Specificity: The exposure leads to a particular outcome rather than a broad grab bag of effects.
  • Experiment: When experimental evidence exists, even from animal models or natural experiments, it supports the association.
  • Analogy: A similar cause-and-effect relationship has been established elsewhere.

These viewpoints have been revisited and debated extensively in the decades since. Some researchers have argued they need updating to incorporate modern causal thinking, particularly advances in computational methods and molecular biology.6PubMed Central. Assessing causality in epidemiology: revisiting Bradford Hill to incorporate developments in causal thinking Others have proposed modernized versions that make the criteria more precise for current data landscapes.7PubMed. Modernizing the Bradford Hill criteria for assessing causal relationships in observational data Still, the core insight endures: no single criterion is sufficient, but when several point in the same direction, the causal interpretation gains weight.

Prospective Studies and the Importance of Timing

Among observational study designs, prospective cohort studies occupy a special niche because they establish something randomized trials also provide: temporal ordering. In a prospective cohort, researchers identify a group of people, measure their exposures or risk factors, and then follow them forward in time to see who develops the outcome of interest. Because the exposure is measured before the outcome occurs, you can be confident you are not mistaking effect for cause.8ScienceDirect. Prospective Cohort Study

The Framingham Heart Study, which has followed residents of a Massachusetts town since 1948, is a classic example. By tracking thousands of people before they developed heart disease, researchers identified smoking, high blood pressure, and high cholesterol as causes of cardiovascular events rather than consequences of them. Without that forward-looking design, the direction of causality would have been much harder to pin down.

Prospective studies do not eliminate confounding the way randomization does, though. The people who choose to smoke differ from nonsmokers in many ways beyond their tobacco use, and not all of those differences can be measured and accounted for. This is why even the best observational study, no matter how large or long-running, sits a step below a well-designed randomized trial in the hierarchy of causal evidence.

Quasi-Experimental Methods

Between the clean causal inference of a randomized trial and the messiness of a standard observational study lies a family of designs that exploit natural or policy-driven variation to approximate an experiment. These quasi-experimental methods have become central to causal research in economics, public health, and social science.

Difference-in-Differences

When a policy change affects one group but not another, researchers can compare how outcomes changed over time in the affected group versus a similar unaffected group. The key assumption is that, without the policy, both groups would have followed the same trajectory. If their outcome trends were parallel before the policy hit and then diverged afterward, the divergence is attributed to the policy.9PubMed Central. Advances in Difference-in-differences Methods for Policy Evaluation Research This approach was famously used to study the employment effects of minimum wage increases by comparing neighboring counties across a state border where only one side raised its minimum wage.

Mendelian Randomization

Genetics offers a clever workaround to the problem of confounding. Because genetic variants are assigned at conception and are not influenced by lifestyle, socioeconomic status, or disease, they can serve as a kind of natural experiment. If a genetic variant that raises your cholesterol level also raises your risk of heart disease, that is strong evidence that the cholesterol itself is doing the damage rather than some confounding behavior. This approach uses genetic variants as instruments to estimate the causal effect of an exposure on an outcome.10PubMed Central. A review of instrumental variable estimators for Mendelian randomization It has become a powerful way to test causal hypotheses that would be unethical or impractical to study in a randomized trial.11American Journal of Epidemiology. Multivariable Mendelian Randomization: The Use of Pleiotropic Genetic Variants to Estimate Causal Effects

Mendelian randomization is not bulletproof, however. It rests on the assumption that the genetic variant affects the outcome only through the exposure of interest and has no other pathways. When a gene influences multiple traits, that assumption can break down. Studies focusing on elderly populations face an additional problem: people who were most susceptible to the exposure may have already died before the study began, warping the results.12PubMed Central. Survival Bias in Mendelian Randomization Studies: A Threat to Causal Inference13Biostatistics. Survivor bias in Mendelian randomization analysis

Regression Discontinuity

Sometimes an intervention is assigned based on whether someone falls above or below a specific cutoff on a continuous score: a test result, an age threshold, or a pollution reading. People just above and just below that line are essentially identical in every way except that one group received the intervention. By comparing outcomes for people near the cutoff, researchers can estimate the causal effect of the intervention as if it had been randomly assigned in that narrow band.14PubMed Central. Regression discontinuity designs in epidemiology: causal inference without randomized trials This design has been used to study everything from the effects of class size on student achievement (using enrollment cutoffs that trigger the addition of a new classroom) to the health effects of medications prescribed above a particular lab value.

Mapping Causes With Directed Acyclic Graphs

Before collecting data or running a statistical model, researchers increasingly draw diagrams of what they believe the causal structure looks like. These diagrams, called directed acyclic graphs, use arrows between variables to represent causal relationships. An arrow from A to B means A is believed to cause B. The “acyclic” part means no loops: you cannot follow the arrows from A back to A.

The value of these diagrams goes beyond illustration. Once you draw out the causal structure, a set of simple rules tells you which variables you need to account for and which you should leave alone in your analysis.15PubMed Central. Directed acyclic graphs for clinical research: a tutorial This matters because adjusting for the wrong variable can introduce bias rather than remove it. A directed acyclic graph helps identify variables that, if controlled for in the design or analysis, are sufficient to eliminate confounding.16PubMed Central. Tutorial on directed acyclic graphs

Consider a study asking whether obesity causes diabetes. You might think controlling for diet quality is helpful. But if obesity causes changes in diet (perhaps through insulin resistance altering cravings), then diet sits on the causal pathway between exposure and outcome. Adjusting for it would block part of the very effect you are trying to measure. A properly drawn causal diagram makes this mistake visible before it happens.

Traps That Fool Even Careful Researchers

Even sophisticated study designs can produce misleading results when certain structural problems go unrecognized. Three of the most common are worth understanding because they undermine causal claims in ways that are not obvious from the data alone.

Collider Bias

A collider is a variable that is caused by two other variables. Conditioning on it, whether by restricting your study population or including it in a statistical model, can create a false association between those two causes. A striking real-world example: among people with diabetes, obesity appears to be associated with lower mortality, even though in the general population the association runs the other way. This misleading finding arises because restricting the analysis to people with diabetes amounts to conditioning on a variable (diabetes) that is influenced by both obesity and other risk factors.17PubMed Central. Collider Bias in Observational Studies The lesson is counterintuitive: adjusting for more variables is not always better. Sometimes it introduces bias that was not there before.

Simpson’s Paradox

A trend that appears in combined data can reverse or disappear when the data is split into subgroups. A treatment might look effective overall but harmful in every age group, or vice versa. The direction of an association at the population level can be reversed within the subgroups that make up that population.18PubMed Central. Simpson’s paradox in psychological science: a practical guide This happens because of differences in how subgroups are distributed across comparison conditions. A hospital that treats sicker patients might have worse outcomes overall even though it performs better within every severity category, simply because its patient mix is more severe.

Reverse Causality and Feedback Loops

When X and Y influence each other over time, separating cause from effect becomes especially tricky. Does poor sleep cause anxiety, or does anxiety cause poor sleep? Often the honest answer is both, creating a feedback loop that standard one-directional methods struggle to untangle. Researchers distinguish between situations where the feedback itself is the phenomenon of interest and situations where it is a nuisance that threatens identification of a causal effect.2Journal of Developmental and Life-Course Criminology. Reciprocal Relationships, Reverse Causality, and Temporal Ordering: Testing Theories with Cross-lagged Panel Models In either case, getting the temporal ordering right and having data measured at enough time points is essential.

The Counterfactual Way of Thinking

Underlying many modern causal methods is a deceptively simple question: what would have happened if things had been different? If a patient received the drug, the causal effect is the difference between their actual outcome and the outcome they would have had if they had not received it. The problem, of course, is that you can never observe both. You see what happened; you never see what would have happened under a different scenario. This counterfactual framework defines individual-level causal effects by comparing what would have been observed under exposure to what would have been observed under no exposure.19Oxford Academic (International Journal of Epidemiology). Commentary: On Causes, Causal Inference, and Potential Outcomes

Randomized trials address this by creating groups that are, on average, identical. The control group stands in for the counterfactual. Observational methods try to approximate the same logic through matching, weighting, or the quasi-experimental designs described above. Target trial emulation, a relatively recent approach, structures the analysis of observational data to mimic what a hypothetical randomized trial would have looked like, reducing bias and making the causal reasoning more transparent.20PubMed. Target Trial Emulation: A Concept Simply Explained

Evidence Triangulation

No single study, no matter how well designed, proves causation beyond all doubt. Every study design has blind spots. Randomized trials can lack real-world relevance. Mendelian randomization can be undermined by genetic variants that affect more than one pathway. Prospective cohort studies cannot eliminate unmeasured confounding. The strongest causal arguments emerge when multiple study designs, each with different vulnerabilities, point to the same conclusion.

This principle has been formalized as evidence triangulation. The idea is to identify the most important weaknesses of any given study approach, then seek out new evidence from approaches that do not share those weaknesses. When results are consistent across studies that rest on different assumptions and for which biases should be unrelated, the conclusion stands on much sturdier ground.21PubMed Central. Evidence triangulation in health research The causal link between smoking and lung cancer, for instance, was built not from a single decisive experiment but from animal studies, prospective cohorts, dose-response relationships, biological mechanisms, and eventually the observation that quitting reduced risk. Each strand was imperfect; together, they were overwhelming.

When Algorithms Try to Find Causes

Machine learning is increasingly being applied to the problem of causal discovery, the task of inferring causal structure directly from data without pre-specifying what you expect to find. Traditional methods rely on researchers drawing the causal diagram first; newer computational approaches attempt to learn the diagram from the patterns in the data themselves. Some methods use supervised machine learning to map observational data to possible causal models.22Journal of Data Science. Causal Discovery for Observational Sciences Using Supervised Machine Learning Others combine machine learning predictions with interpretability techniques to identify and rank the strength of causal relationships among variables.23PubMed Central. ReX: Causal Discovery based on Machine Learning and Explainability techniques

The appeal is obvious: in fields where thousands of variables interact, no human can draw every plausible causal diagram by hand. But the limitations are equally real. Machine learning excels at prediction, which is not the same thing as causation. A model might learn that umbrella sales predict rain, but only because both are linked to weather forecasts. Without domain knowledge to interpret the output, algorithmic causal discovery can generate plausible-looking but meaningless arrows. These tools are best understood as hypothesis generators: they can suggest causal structures that researchers then test with the more rigorous methods above, rather than final answers in themselves.

P-hacking and the Fragility of Published Causal Claims

Even when the right method is used, the incentive structure of scientific publishing can distort the evidence base. Journals overwhelmingly publish positive results, creating pressure to find statistical significance. One widespread problem is that researchers, sometimes unconsciously, keep adjusting how they collect, select, or analyze data until a nonsignificant result crosses the threshold for significance. Text-mining across the scientific literature has demonstrated that this practice is widespread.24PubMed Central. The extent and consequences of p-hacking in science

This problem is not confined to simple correlational studies. An analysis of over 7,500 hypothesis tests from causal study designs published in leading journals found systematic clustering of results just above the conventional significance threshold, suggesting that publication incentives and researcher flexibility may distort even the evidence from difference-in-differences analyses, instrumental variable studies, and randomized controlled trials.25Information Systems Research. p-Hacking and Publication Bias in Design-Based Causal Studies in Information Systems A separate large-scale examination of economics journals found that the extent of the problem varies by method.26American Economic Review. Methods Matter: p-Hacking and Publication Bias in Causal Analysis in Economics The practical implication is sobering: a single published study claiming X causes Y, even one using a rigorous design, might represent the one analysis out of many that happened to produce a significant result. This is another reason evidence triangulation matters so much. A causal claim supported by only one study in one field using one method deserves skepticism, regardless of how elegant the method is.

Time Series and Granger Causality

In fields where data arrives as sequences over time, such as neuroscience, economics, and climate science, a different approach to causation has taken hold. Rather than asking whether X causes Y in the philosophical sense, Granger causality asks a narrower question: does knowing the past values of X improve your ability to predict Y, beyond what Y’s own past values tell you? If so, X is said to “Granger-cause” Y. These methods were developed to analyze the flow of information between time series.27PubMed Central. A study of problems encountered in Granger causality analysis from a neuroscience perspective

The name is slightly misleading. Granger causality is really about predictive precedence, not causation in the way most people mean the word. If stock prices in one market consistently move before prices in another, the first “Granger-causes” the second, but this could reflect a shared driver hitting one market sooner. Still, in fields like neuroscience, where researchers want to know whether one brain region’s activity drives another’s, the framework provides a useful and testable starting point, as long as its limitations are kept in mind.