What Are the 5 Steps of the Scientific Method?

The five steps of the scientific method, as typically taught, are: (1) make an observation and ask a question, (2) form a hypothesis, (3) test the hypothesis through experimentation, (4) analyze the data, and (5) draw a conclusion. This tidy sequence has anchored science education for generations, and it does capture something real about how knowledge gets built. But practicing scientists will tell you that the actual process is far messier, more iterative, and more creative than any five-step recipe suggests.

Observation and the Question That Starts Everything

Every investigation begins with noticing something. Sometimes the observation is deliberate: a researcher carefully measuring rainfall chemistry over months, for instance. Other times it arrives uninvited. The discovery of acid rain is a good example. While the English chemist Robert Angus Smith first coined the term in the late 1800s after systematically studying rain composition, the phenomenon was independently rediscovered by scientists who weren’t looking for it at all. Hub Vogelman, a botanist at the University of Vermont, was gathering precipitation data on Camel’s Hump mountain when he noticed something off in his numbers and connected it to the large numbers of dying spruce trees in the area. He hadn’t set out to find acid rain; careful, systematic observation put him in a position to recognize it anyway.1ILAR Journal. Observation and Cogitation: How Serendipity Provides the Building Blocks of Scientific Discovery

The observation stage blurs into question-asking almost immediately. You notice dying trees and acidic rainfall, and the question writes itself: is the rain killing the trees? This is the seed of a research question, and it shapes everything that follows. A vague observation without a focused question leads nowhere. A sharp question pointed at the wrong phenomenon wastes years. The quality of the initial observation, and the curiosity it triggers, matters enormously.

Forming a Hypothesis

A hypothesis is a proposed explanation for what you observed, stated in a way that you can actually test. “Acidic rain is damaging spruce trees at high elevation” is testable. “Nature is angry” is not. The key feature of a good hypothesis is that it makes a prediction you can check against reality, and it risks being wrong. The philosopher Karl Popper made this the centerpiece of his thinking about science: a claim only counts as scientific if it is falsifiable, meaning some conceivable observation could prove it wrong.2PubMed Central. Falsifiability in medicine: what clinicians can learn from Karl Popper Popper argued that simply piling up confirming examples for a theory isn’t enough. The theory has to stick its neck out and make bold predictions that risk refutation.

In practice, hypotheses rarely emerge from thin air. They grow out of previous work, existing data, and informed guesses about how something might operate. Research on hypothesis formulation emphasizes that a testable working hypothesis grounded in earlier evidence is the starting point for original research, and that hypotheses lacking evidence-based justification tend to be poorly received by the scientific community.3PubMed Central. Formulating Hypotheses for Different Study Designs You don’t just guess randomly. You read what’s already been found, identify a gap or a tension, and propose something that could fill it.

Testing Through Experimentation

Once you have a testable prediction, you design an experiment to see whether reality cooperates. The hallmark of a good experiment is control: you change one thing (the variable you care about), hold everything else constant, and compare what happens to a group that didn’t get the change. Experimental controls exist because human senses and intuitions are unreliable. Without a control group, you can’t tell whether what you observed was caused by your variable or by something else entirely.4PubMed Central. Why control an experiment?: From empiricism, via consciousness, toward Implicate Order

Not all sciences get to run experiments in the classic sense. Astronomers can’t manipulate stars. Paleontologists can’t rerun extinction events. These fields rely on careful observation of events that happened or are happening without human intervention, rather than actively manipulating conditions.5Oxford Academic. Experiment, observation and the confirmation of laws The distinction between experimental and observational science is important because it changes what kinds of conclusions you can draw. An experiment can point toward causation. An observation, no matter how detailed, usually can only show correlation. Both are legitimate, but they carry different evidentiary weight.

Analyzing the Data

After you collect your results, you have to figure out what they mean. This is where statistical tools come in. Researchers use hypothesis testing to ask a simple question: could the results I got have happened by chance alone? The p-value, for instance, measures how likely it would be to see results as extreme as yours if there were actually no real effect. It functions as a decision-making tool rooted in probability.6PubMed Central. The nuts and bolts of hypothesis testing

Analysis sounds mechanical, but it’s where human judgment and human fallibility matter most. The data don’t speak for themselves. You choose which patterns to look for, which comparisons to highlight, and how to handle messy or ambiguous results. Those choices shape the story you end up telling, and they’re susceptible to all sorts of conscious and unconscious biases.

Drawing Conclusions and Sharing Them

The final step in the textbook version is drawing a conclusion: does the evidence support or refute your hypothesis? If the data line up with your prediction, the hypothesis survives for now, though it hasn’t been “proved.” If the data contradict it, you revise or discard the hypothesis and try again. Science moves forward through this cycle of proposing, testing, and revising.

Conclusions don’t just sit in a notebook. Scientists write up their findings and submit them for peer review, a process in which other experts evaluate the work for quality before it’s published.7PubMed Central. Peer review guidance: a primer for researchers Peer review isn’t perfect, but it serves as a filter. Reviewers look for problems with the experimental design, the statistical analysis, and the logic connecting evidence to conclusions. If a study survives that scrutiny and gets published, other scientists can attempt to replicate it, building confidence in the finding over time or revealing that it doesn’t hold up.

Why Five Steps Is a Simplification

The five-step version you find in textbooks is a useful teaching scaffold, not a description of how science actually works. One detailed account of the scientific method lists not five but eight distinct steps: making observations, incorporating background research, generating a hypothesis, designing a controlled experiment, collecting data, revising the hypothesis and collecting more data in an iterative loop, analyzing results, and communicating findings.8Academic Press. Scientific method The iteration part is the critical addition. In practice, you don’t march from step one to step five in a straight line and call it done. You often loop back: an unexpected result sends you to a new hypothesis, a better experiment, a revised analysis. The path from question to answer is rarely linear.

Science education researchers have raised concerns that teaching the method as a rigid sequence of discrete steps can actually distract students from productive inquiry.9Science Education. The scientific method and scientific inquiry: Tensions in teaching and learning When students (or adults) internalize the method as a fixed checklist, they tend to focus on procedure rather than understanding. The goal of science is to figure out how the world works, not to perform a ritual in the right order. A related critique is that the standard classroom version emphasizes testing predictions at the expense of deeper conceptual thinking, producing graduates who can follow a protocol but struggle to explain why the protocol matters.10Science Education. Beyond the scientific method: Model‐based inquiry as a new paradigm of preference for school science investigations

When Whole Discoveries Skip Steps

Here is something genuinely surprising: a significant fraction of major scientific discoveries didn’t follow the standard method at all. A study examining discoveries since 1900 found that roughly a quarter of them did not apply all three core features of the traditional method (observation, experimentation, and hypothesis testing). About 6% involved no formal observation, 23% used no experimentation, and 17% never tested a hypothesis.11PubMed Central. Redefining the scientific method: As the use of sophisticated scientific methods that extend our mind Some of the most important advances in mathematics, theoretical physics, and computer science were built entirely from logical deduction and modeling rather than from experiments. Darwin’s theory of evolution was assembled from years of observational evidence, not controlled experiments in the traditional sense.

This doesn’t mean the scientific method is useless. It means the five-step version captures one common path to knowledge, not the only one. Fields like astronomy, paleontology, and epidemiology have always relied heavily on observational data because you can’t run controlled experiments on stars, fossils, or pandemics. The underlying logic of science, making claims that can be tested against evidence and revised when they’re wrong, still applies. But the specific sequence of steps varies enormously depending on what you’re studying.

How Confirmation Bias Undermines the Process

One of the biggest threats to good science is baked into human psychology. People tend to seek out evidence that supports what they already believe and to overlook evidence that contradicts it. This is confirmation bias, and it doesn’t spare trained scientists. In a simulated research environment, subjects showed a strong tendency to choose settings that would confirm their existing hypotheses rather than settings that could test alternative explanations. The encouraging finding was that when people did encounter clear falsifying information, they used it to reject incorrect ideas, but they had to be forced into the position of seeing it first.12Quarterly Journal of Experimental Psychology. Confirmation Bias in a Simulated Research Environment: An Experimental Study of Scientific Inference

More recent work on how people gather evidence paints a similar picture. When people are confident in an initial choice, they actively sample new evidence in ways that favor that choice, creating a self-reinforcing loop.13PubMed Central. Humans actively sample evidence to support prior beliefs For scientists, this means the analysis and conclusion steps of the method are vulnerable. A researcher who expects their hypothesis to be confirmed may unconsciously design analyses, choose subgroups, or interpret ambiguous results in ways that make confirmation more likely. The method’s emphasis on falsifiability is partly meant to counteract this tendency, but awareness of the bias is the first line of defense.

P-Hacking and the Replication Crisis

Confirmation bias operating at scale helped produce what researchers now call the replication crisis: the discovery that many published scientific findings can’t be reproduced when other teams try. One major contributing factor is a set of practices collectively known as p-hacking, essentially massaging data or analysis methods until a result crosses the threshold of statistical significance. A comprehensive review identified 12 distinct p-hacking strategies and demonstrated through simulations that these strategies dramatically inflate false-positive rates.14PubMed Central. Big little lies: a compendium and simulation of p-hacking strategies

Other questionable practices that undermine the method’s integrity include HARKing (hypothesizing after results are known), which flips the sequence on its head. Instead of predicting an outcome and testing it, a researcher finds an interesting pattern in the data and then writes the paper as though that pattern was predicted all along. The proliferation of these practices has driven efforts at reform, including pre-registration of studies, where researchers publicly commit to their hypothesis and analysis plan before collecting data.15PubMed Central. Campbell’s Law Explains the Replication Crisis: Pre-Registration Badges Are History Repeating Pre-registration doesn’t prevent all manipulation, but it makes it much harder to quietly rearrange your hypothesis after peeking at the results.

The replication crisis is uncomfortable, but it’s also the scientific method working as intended at a larger scale. Scientists noticed a pattern (too many findings failing replication), proposed an explanation (questionable research practices), tested it (by examining published methods and running replication studies), and are now revising their procedures. The self-correcting nature of science is real, even if it operates on a timescale of decades rather than days.

Ethics Review as an Unofficial Step

One stage that the five-step model never mentions is ethics review. Before any experiment involving human participants (and most involving animals) can begin, the research plan must be approved by an ethics committee. These committees evaluate whether the study design is safe, whether participants are adequately informed and protected, and whether the potential benefits justify any risks.16PubMed Central. Ethics Committees: Structure, Roles, and Issues In biomedical research, this step is legally required, not optional. It happens between hypothesis formation and data collection, and it can reshape or even halt a study. The textbook version leaves this out entirely, but for anyone conducting research with living subjects, ethics review is as fundamental as the experiment itself.

Data-Driven Science and the Fourth Paradigm

The traditional scientific method is hypothesis-driven: you start with an idea and design a test. But an increasingly large share of modern research works in the opposite direction. In data-driven science, researchers start with enormous datasets and use computational tools to find patterns that no human would have thought to look for. Machine learning algorithms can sift through millions of data points, identifying relationships between variables that weren’t part of any hypothesis. This approach has been described as the “fourth paradigm” of scientific discovery, following experimental science, theoretical science, and computational simulation.17PubMed Central. A Comparison of Hypothesis-Driven and Data-Driven Research: A Case Study in Multimodal Data Science in Gut-Brain Axis Research

Data-driven research doesn’t abandon the scientific method so much as rearrange it. The patterns discovered through data mining still need to be validated. A machine-learning algorithm might flag a correlation between gut bacteria and brain activity, but that finding only becomes scientifically meaningful when someone forms a hypothesis about why the correlation exists and tests it with an independent dataset or a controlled experiment. The five-step model struggles to capture this kind of workflow, where exploration precedes hypothesis rather than following observation. But the underlying logic of evidence, testing, and revision still applies.

Why Scientists Emphasize What They Don’t Know

If you’ve ever been frustrated by a scientist saying “we think” or “the evidence suggests” instead of giving a straight answer, there’s a reason for it. Uncertainty isn’t a weakness in science; it’s the honest reporting of where the evidence stands. Research on public communication during the COVID-19 pandemic found that downplaying uncertainty could raise short-term support for scientific recommendations, but that when projections later turned out to be wrong, public trust in science took a hit.18PubMed Central. Model uncertainty, political contestation, and public trust in science: Evidence from the COVID-19 pandemic Stating clearly what is known, what is uncertain, and what could change turns out to be better for long-term credibility than projecting false confidence.

The five-step model can inadvertently make this worse by implying that science produces clean, final answers. A student who learns “observe → hypothesize → experiment → analyze → conclude” might reasonably expect the conclusion to be the end of the story. In reality, every conclusion opens new questions, and most findings come with error bars, caveats, and conditions. The method is a loop, not a conveyor belt, and the conclusions it produces are the best answers available right now, subject to revision as better evidence arrives. Understanding this iterative character is more useful than memorizing the steps themselves.