Error bars are the thin lines extending above and below data points or bar tops on a graph, and they represent some measure of variability or uncertainty in the data. The catch is that “some measure” can mean very different things depending on which type of error bar the graph uses. An error bar showing a standard deviation, a standard error, or a 95% confidence interval each tells you something fundamentally different, yet they can all look identical on a graph. That ambiguity is the source of most confusion about error bars, and it is exactly what makes knowing the type so important before drawing any conclusions.
Three Types of Error Bars and What Each One Means
Nearly every error bar you encounter in a scientific graph falls into one of three categories: standard deviation, standard error of the mean, or confidence interval. They are calculated differently, they answer different questions, and reading one as if it were another leads to wrong conclusions.
Standard deviation (SD) bars describe how spread out the individual data points are. If a graph shows the average height of a group of people with SD error bars, those bars tell you the range within which most individual heights fall. They are about the population’s variability, not about how confident you should be in the average itself. As one analysis puts it, SD bars “only tell us about the spread of the population” and “do not technically indicate anything about the uncertainty of a mean or statistical significance.”1Nature Methods. How to Properly Interpret Error Bars A researcher measuring blood pressure in 50 patients would use SD bars to show the reader how much blood pressure varied from person to person.
Standard error of the mean (SEM) bars are smaller than SD bars, often much smaller, because they answer a narrower question: how precisely has this experiment estimated the true average? The SEM shrinks as the sample size grows. Measure 10 people and the SEM might be large; measure 10,000 and it becomes tiny, even if the individual variation (SD) is exactly the same. The relationship is straightforward: the SEM equals the SD divided by the square root of the sample size.2PubMed Central. A note on error bars as a graphical representation of the variability of data in biomedical research: Choosing between standard deviation and standard error of the mean This means SEM bars always look more impressive on a graph. They make data appear tighter and more precise, which is one reason some researchers prefer them, and one reason readers need to check the label.
Confidence interval (CI) bars, typically set at 95%, give you a range within which the true population mean is likely to fall. They are the most directly useful for judging whether two groups differ from each other, because they combine information about both variability and sample size into a single statement about certainty. A 95% CI says: if we repeated this experiment many times, about 95% of those intervals would contain the true value.
Why You Cannot Interpret Error Bars Without Reading the Label
Here is the core problem: a graph with SD bars, a graph with SEM bars, and a graph with 95% CI bars can look nearly identical to the eye. The bars extend, the data points sit at the same heights, and nothing about the visual format tells you which type you are looking at. The only way to know is to read the figure legend or axis label. If the graph does not specify what the error bars represent, it is essentially uninterpretable.
This is not just an abstract concern. Different types of error bars “give quite different information, and so figure legends must make clear what error bars represent.”3PubMed Central. Error bars in experimental biology Yet published papers frequently omit this information or bury it in footnotes. One review of articles from high-impact scientific journals found that investigators across fields were inconsistent in which type of error bar they presented, with no standard practice even within a single journal.2PubMed Central. A note on error bars as a graphical representation of the variability of data in biomedical research: Choosing between standard deviation and standard error of the mean
It is also worth noting that standard deviation bars are sometimes labeled as “error bars” even though some researchers argue they should not be called that at all. SD bars describe sample variability around the sample mean rather than serving as a precision measure for estimating the true population mean.4Rev. Soc. Bras. Med. Trop.. Description of continuous data using bar graphs: a misleading approach They are not “errors” in the sense of uncertainty about the average; they are descriptions of natural variation. Calling them error bars conflates two very different ideas, which feeds the confusion.
Judging Whether Two Groups Differ by Looking at Error Bars
The most common reason people look at error bars is to decide whether two groups are really different or whether the apparent gap could just be noise. Researchers and readers alike tend to eyeball the overlap between the bars on two data points and draw conclusions. This works, but only if you know the rules, and the rules change depending on the bar type.
For 95% confidence intervals: if the bars of two groups do not overlap at all, the difference between those groups is statistically significant. If they overlap moderately, the difference might still be significant. CIs have to overlap by roughly half their length before you can be confident the difference is not significant. Many people assume any overlap at all means no difference, but that is too conservative.
For standard error bars: the gap has to be larger. Two SEM bars need to be separated by a gap of roughly one full bar length (not just touching) for the difference to reach roughly a 95% significance level when comparing two independent groups. This is because SEM bars are narrower than CI bars to begin with, so a given visual gap represents less statistical distance.
For standard deviation bars: overlap tells you almost nothing about whether the groups are statistically different. SD bars reflect the spread of individual measurements, not the precision of the group average. Two groups can have massively overlapping SD bars and still have a highly significant difference between their means, especially if the sample sizes are large. Trying to judge significance by eyeballing SD overlap is a common and serious mistake.
These visual rules are useful shortcuts, not replacements for actual statistical tests. But they highlight why confusing one bar type for another can flip your conclusion entirely. If you treat SEM bars as though they were CI bars, you will underestimate uncertainty. If you treat SD bars as though they indicate precision, you will see differences that are not there, or miss ones that are.
Even Experts Misread Error Bars
If you find error bars confusing, you are in good company. In one study, about a third of peer-reviewed authors who were surveyed misinterpreted error bars showing confidence intervals or standard error when judging significance from a graph.5PLoS One. Revealing undergraduate biology students’ conception of variability and error bars within graphing These were not students or casual readers; they were researchers who publish in scientific journals. The error rate was 31.5%, meaning nearly one in three active scientists drew the wrong conclusion from a graph containing error bars.
Among students, the picture is even bleaker. Research on undergraduate biology students found that many struggled to recognize whether variability was even present in a graph when asked to interpret data within a treatment group. Students who had only graphed raw data on bar charts showed less understanding of error bars than those who had experience graphing means with error bars attached.5PLoS One. Revealing undergraduate biology students’ conception of variability and error bars within graphing In other words, the ability to read error bars is not intuitive. It is a learned skill, and formal education does not always teach it well.
This gap matters because error bars appear everywhere: in medical research, climate data, drug trial results, public health reports, and news graphics. When a third of the scientists producing those graphs misread them, the downstream consequences for everyone who encounters them secondhand are significant.
The Within-the-Bar Bias
Beyond misinterpreting the type of error bar, people have a deeper cognitive quirk when reading bar graphs that most never notice. Researchers have documented what they call the “within-the-bar bias”: when viewers see a bar depicting a mean value and are asked to judge whether a particular data point is likely to come from the same distribution, they consistently judge points that fall inside the shaded area of the bar as more likely than points the same distance from the mean but outside the bar.6PubMed. Bar graphs depicting averages are perceptually misinterpreted: the within-the-bar bias
Think about what that means. A bar graph showing that the average test score was 75 has a bar that stretches from the x-axis up to 75. A score of 70 (five points below the mean, inside the bar’s shaded area) and a score of 80 (five points above the mean, outside the bar) are equally likely to appear in the data. But viewers consistently rate the 70 as more plausible than the 80, as if the bar somehow “contains” the data. This bias showed up across multiple experiments, persisted whether error bars were present or not, appeared both in memory and in real-time perception, and occurred in college students, community adults, and online samples alike.6PubMed. Bar graphs depicting averages are perceptually misinterpreted: the within-the-bar bias
The practical takeaway is that bar graphs with error bars can create a false sense of containment. You see the bar plus its whiskers and unconsciously treat that shaded-plus-whisker region as “where the data lives,” which is not quite right for any of the three bar types. Being aware of the bias helps, but some data-visualization experts have argued that dot plots or box plots communicate uncertainty more honestly than bars, partly because they avoid triggering this perceptual trap.
When Standard Error Bars Can Mislead
Standard approaches to error bars assume that each data point comes from an independent source: different people, different samples, different experiments. But many studies use within-subjects (repeated-measures) designs, where the same participants are measured multiple times under different conditions. In those cases, regular between-subjects error bars can be actively misleading.
Here is why. Imagine testing whether people react faster to red lights than to green lights. You test the same 20 people on both colors. Each person has their own baseline speed: some are generally fast, others slow. What you care about is whether each person is faster on red than green, not whether the group averages look different when you ignore that each pair of measurements came from the same person. Standard SEM or CI bars calculated the usual way include all that between-person variability, which inflates the bars and makes it harder to see the within-person effect.
Because of the correlational structure in repeated-measures designs, the calculation and interpretation of confidence intervals becomes nontrivial.7PubMed Central. Standard errors and confidence intervals in within-subjects designs: generalizing Loftus and Masson (1994) and avoiding the biases of alternative accounts Specialized methods exist to compute adjusted error bars that account for the paired nature of the data.8Advances in Methods and Practices in Psychological Science. Summary Plots With Adjusted Error Bars: The superb Framework With an Implementation in R These adjusted bars are typically smaller and give a more honest picture of the precision of the within-subjects comparison. But many published graphs in psychology and neuroscience still use unadjusted bars, which means the error bars you see may not match the statistical test the authors actually ran. The visual impression and the reported p-value can point in opposite directions.
If you are reading a paper that uses a within-subjects design, check whether the figure legend mentions adjusted or within-subjects error bars. If it does not, be cautious about drawing conclusions from the visual overlap alone. The statistical test in the text is a better guide than the graph.
How SEM Bars Can Make Data Look Better Than It Is
Because SEM bars are always smaller than SD bars from the same dataset, choosing to display SEM instead of SD makes the data appear less variable and the group differences appear cleaner. This is not dishonest per se. SEM answers a legitimate question about precision. But when graphs show bars that barely peek above the data points and create an impression of rock-solid certainty, it is worth checking whether those bars are SEM rather than SD, and whether the sample size is large enough that the SEM compression is warranted rather than cosmetic.
A dataset with enormous individual variation (say, highly variable patient responses to a drug) can produce tiny SEM bars if the study enrolled enough people. Those tiny bars correctly convey that the average is precisely estimated, but they hide the fact that individual responses were all over the map. For a clinician trying to predict what will happen to a single patient, the SD is the more useful number. For a researcher asking whether the drug shifts the average, SEM or CI is more informative. Neither is wrong; they answer different questions. The problem arises when graphs display one while the reader assumes the other.
A good rule of thumb: if you see remarkably small error bars and the study reports a large sample, those are probably SEM bars. Look at the legend. If the legend says SEM or SE, mentally scale the bars up by multiplying their length by the square root of the sample size, and you will get a rough sense of the SD, which tells you how variable the individual data points actually were.
Asymmetric Error Bars and Logarithmic Scales
Most error bars extend equally above and below the mean, but not always. When data are plotted on a logarithmic scale, the same plus-or-minus range that looks symmetric on a linear axis becomes asymmetric: the upward bar is shorter than the downward bar, or vice versa. This happens because equal distances on a log scale correspond to multiplicative rather than additive changes. A researcher plotting bacterial growth rates or chemical concentrations on a log axis might show error bars that look lopsided, which is not an error in the graph but a natural consequence of the scale.9PLOS ONE. Problems with Using the Normal Distribution – and Ways to Improve Quality and Efficiency of Data Analysis
Asymmetric bars also show up when the data are skewed or when the analysis uses a method that produces asymmetric confidence intervals, such as for proportions near 0% or 100%. If a study reports that 3% of patients had a side effect, the error bar cannot extend equally in both directions because it cannot go below zero. The downward bar gets clipped, producing an asymmetric appearance. Whenever you encounter lopsided error bars, check whether the axis is logarithmic or whether the quantity being measured has a natural boundary. In either case, the asymmetry is usually a feature of honest reporting, not a sign that something went wrong.
No Universal Standard Across Fields
One of the most frustrating aspects of error bars is that different scientific disciplines have different customs about which type to use, and none of those customs are universally enforced. In biomedical research, SEM bars are extremely common, partly because they are smaller and make graphs look cleaner, and partly because clinical researchers are often most interested in how precisely a treatment effect is estimated. In ecology and evolutionary biology, SD bars are more traditional because the natural variability of organisms is itself a finding worth communicating. In physics and engineering, error bars often represent measurement uncertainty, which is a different concept entirely: how precisely the instrument can measure, rather than how variable the thing being measured is.
A review of articles from representative high-impact journals found that investigators remain uncertain about which type of error bar to present, underscoring the lack of a universal standard in the scientific community.2PubMed Central. A note on error bars as a graphical representation of the variability of data in biomedical research: Choosing between standard deviation and standard error of the mean Some journals now require authors to specify the error bar type in their figure guidelines, but compliance is inconsistent. As a reader, the safest approach is to treat every graph’s error bars as ambiguous until you confirm the type from the legend.
A Practical Checklist for Reading Any Graph With Error Bars
When you encounter error bars in a graph, a few quick checks can save you from drawing false conclusions:
- Check the legend: Find out whether the bars represent SD, SEM, or CI. If the legend does not say, treat the graph with skepticism. You literally cannot interpret the bars without this information.
- Check the sample size: SEM bars shrink as sample size grows, even when the underlying variability stays the same. A tiny SEM bar from a study of 5,000 people tells a different story than the same bar from a study of 12.
- Match the bar type to your question: If you want to know how variable the data are (how different individual measurements were from each other), you need SD bars. If you want to know how precisely the average is estimated, you need SEM or CI bars. If you want to judge whether two groups differ, CI bars give you the most direct visual answer.
- Do not assume overlap equals no difference: For CI bars, moderate overlap can still mean a significant difference. For SD bars, overlap tells you almost nothing about significance.
- Watch for within-subjects designs: If the same participants were measured under each condition, standard error bars may overstate the uncertainty of the comparison. Look for adjusted or within-subjects error bars in the legend.
None of this requires statistical training. It requires the same kind of label-reading you would do when comparing nutrition facts on two food packages. The information is there; you just have to look for it before letting the visual impression do the thinking for you.
Alternatives to Traditional Error Bars
Given how often error bars are misread, some researchers and data visualization specialists have moved toward alternatives that communicate uncertainty more transparently. Box plots show the median, the interquartile range, and individual outliers all at once, giving a much richer picture of the data’s shape than a single bar and whisker. Violin plots go further by displaying the full distribution, letting you see whether the data are skewed or bimodal rather than forcing everything into a symmetric summary.
Strip plots or dot plots overlay the actual individual data points on the graph, sometimes combined with a summary bar or line. When data points are visible, readers can directly see the spread, the clustering, and the outliers without needing to decode what a bar length means. Studies on the within-the-bar bias suggest this kind of transparency helps counteract the cognitive tendency to treat the shaded region of a bar as a container for the data. If a graph shows every data point and adds a mean line with CI whiskers on top, you get both the raw picture and the inferential summary. That combination communicates far more than error bars alone ever can, and it is increasingly expected by journals in fields like psychology, ecology, and genomics. When you see a graph that shows individual data points alongside summary statistics, you are looking at a more honest representation of the evidence than a bar graph with error bars ever provides.