How to Calculate Error and Uncertainty in Chemistry

Calculating error and uncertainty in chemistry comes down to identifying every source of variability in a measurement, assigning a numerical value to each one, combining them mathematically, and reporting the final result with a range that tells the reader how confident you should be in that number. The process has been formalized internationally through the Guide to the Expression of Uncertainty in Measurement (known as the GUM), and it applies whether you are titrating a solution in an undergraduate lab or certifying a pharmaceutical reference standard. The core idea is simpler than the formal framework makes it look, but real-world application requires attention to sources of error that are easy to overlook.

Random Error Versus Systematic Error

Every measurement in chemistry is affected by two broad categories of error, and handling them requires different strategies. Random error is the scatter you see when you repeat the same measurement multiple times under the same conditions. One reading comes in a little high, the next a little low, and the spread around the average reflects the precision of your method. You can reduce random error by making more measurements and averaging them, and you can quantify it statistically.

Systematic error, on the other hand, is a consistent bias that pushes all your results in the same direction. A balance that reads 0.002 grams too heavy, a spectrophotometer with a drift in its baseline, or an uncorrected temperature effect on a volumetric flask all produce systematic error. No amount of averaging will fix a systematic offset because every measurement is shifted by the same amount. Detecting systematic error usually requires comparing your results against a known reference material or a different method entirely. When estimating total uncertainty, both types must be accounted for, and the GUM framework provides a structure for doing so.

Type A and Type B Uncertainty Evaluation

The GUM divides uncertainty estimation into two approaches based on how you obtain the information. Type A evaluation uses statistical analysis of repeated measurements. If you measure the same sample several times, the standard deviation of those results captures the random scatter, and the standard deviation of the mean (the standard deviation divided by the square root of the number of measurements) gives you the standard uncertainty for that source. The more measurements you make, the smaller that standard uncertainty becomes, because the average gets more reliable even if individual readings still bounce around.

Type B evaluation covers everything you did not measure yourself but can still quantify. This includes tolerances printed on a volumetric flask, calibration certificates for your balance, published reference values for a standard, and manufacturer specifications for an instrument’s linearity or drift. For example, a Class A 100 mL volumetric flask has a printed tolerance, and you convert that tolerance into a standard uncertainty using an assumed probability distribution (usually rectangular if you have no other information, meaning you treat every value within the tolerance as equally likely). Type B sources often outnumber Type A sources in a real uncertainty budget, because most of the equipment in a laboratory comes with documented specifications rather than being repeatedly tested on the spot.

Standard uncertainties from both types are expressed in the same units and are combined in the same way, which is one of the strengths of the GUM approach: it does not matter whether you got the number from statistics or from a certificate, once it is a standard uncertainty it enters the same calculation.

Combining Uncertainties Through Propagation

Most results in chemistry are not a single direct measurement but a calculation that depends on several measured quantities. A concentration from a titration depends on the volume of titrant, the molarity of the titrant, the mass of the sample, and possibly a dilution factor. Each of those quantities carries its own uncertainty, and you need a way to figure out how they combine to affect the final answer.

The classical approach uses what is often called the law of propagation of uncertainty. For quantities that are added or subtracted, you combine their absolute uncertainties by taking the square root of the sum of their squared uncertainties. For quantities that are multiplied or divided, you do the same thing with relative uncertainties (each uncertainty divided by its measured value). The squaring-and-square-rooting step reflects the fact that independent errors are unlikely to all push in the same direction at once; the combined effect is generally smaller than if you simply added every individual uncertainty together.

This propagation step is where the uncertainty budget takes shape. You list every input quantity, assign each one a standard uncertainty (from Type A or Type B evaluation), determine how sensitive the final result is to each input, and combine them. The output is the combined standard uncertainty of your result, a single number that captures the total expected spread from all identified sources.

Expanded Uncertainty and How to Report a Result

A combined standard uncertainty by itself represents roughly a 68 percent confidence interval, assuming the uncertainties follow a normal distribution. In practice, most laboratories and regulatory frameworks want a higher level of confidence, so the combined standard uncertainty is multiplied by a coverage factor to produce an expanded uncertainty. A coverage factor of 2 corresponds to roughly 95 percent confidence, meaning the true value is expected to lie within the reported range about 95 times out of 100. This convention is widely used across analytical chemistry and clinical chemistry alike.

Results are then reported as the measured value plus or minus the expanded uncertainty, along with the coverage factor used. For instance, a calcium content might be reported as 1531 ± 177 mg per kilogram with a coverage factor of 2 at approximately 95 percent confidence. Stating the coverage factor is important because without it, the reader cannot tell whether the reported range represents one standard uncertainty or two, which would change its meaning substantially.

Where Weighing Goes Wrong

Mass measurements seem straightforward, but the uncertainty in a balance reading involves more than just the repeatability of the display. Air buoyancy is a persistent systematic effect: an object weighed in air appears lighter than its true mass because the displaced air exerts an upward force, and the magnitude depends on the density of the object, the density of the calibration weights, and the density of the ambient air. For high-accuracy work, air buoyancy can be a larger source of bias than the balance’s own repeatability. Beyond that, nonlinearity in the balance’s response across its range often contributes more to the combined uncertainty than repeatability alone, with smaller contributions from temperature sensitivity and the calibration data of the balance itself.

In a student laboratory, these effects are usually negligible compared to other sources of error, but in reference material certification or pharmaceutical quality control, ignoring them can lead to uncertainty estimates that are misleadingly small. The practical takeaway is that a balance’s displayed precision (say, ±0.0001 g) is not the full picture of how uncertain your mass value really is.

Calibration Curves and Regression Uncertainty

Instrumental methods in chemistry almost always rely on a calibration curve: you measure the instrument’s response for a set of standards at known concentrations, fit a line through the data, and then use that line to convert an unknown sample’s response into a concentration. The uncertainty in the final concentration comes not just from the noise in the unknown’s signal but also from the uncertainties in the slope and intercept of the regression line.

When you interpolate an unknown from a calibration curve, the confidence interval around the predicted concentration depends on how many calibration points you used, how much scatter they showed, and how far the unknown’s response is from the center of the calibration range. Measurements near the middle of the calibration curve have smaller uncertainty than those near the edges, which is one reason analysts are taught to bracket their samples within the calibration range rather than extrapolating beyond it. In a worked example from the literature, a sample with a signal of 0.030 measured against a five-point calibration yielded a concentration of about 2.84 micrograms per liter with a 95 percent confidence interval of roughly ±0.12 micrograms per liter.

Sample Preparation as a Hidden Source of Error

One area that often gets overlooked in uncertainty budgets is the sample preparation step. Grinding, dissolving, extracting, filtering, and diluting a sample all introduce variability, and this variability can be surprisingly large relative to the measurement itself. Research on physical sample preparation has shown that the preparation process can contribute up to about 20 percent of total variability, with relative uncertainties for individual analytes reaching as high as 66 percent at the 95 percent confidence level in difficult cases.

The reason this matters is that many uncertainty estimates focus almost entirely on the instrument and the calibration, ignoring the fact that the sample sitting in the autosampler may not perfectly represent the original material. Homogeneity of the sample, losses during extraction, and contamination during handling are all real effects. If your uncertainty budget does not include a term for sample preparation, you may be reporting a result that looks more precise than it actually is.

Detection Limits and the Edge of Measurability

At very low concentrations, uncertainty becomes so large relative to the measured value that the result stops being meaningful. This is where detection limits come in. The limit of detection is the lowest concentration at which you can reliably distinguish a signal from background noise. A common rule of thumb defines it as three times the standard deviation of the blank, but more rigorous approaches distinguish between the limit of blank, the limit of detection, and the limit of quantitation.

The limit of blank is the highest apparent concentration you would expect to see when measuring a sample that contains none of the analyte at all. It accounts for the noise floor of the instrument. The limit of detection builds on this by also considering the variability of a sample at a genuinely low concentration, defining the lowest level at which detection is feasible. The limit of quantitation goes a step further, setting the lowest concentration at which you can report a result with acceptable precision, sometimes defined as the concentration where the relative standard deviation drops to about 10 percent.

Below the limit of quantitation, your uncertainty is so large compared to the measured value that reporting a specific number is misleading. Laboratories typically report such results as “less than” the quantitation limit rather than giving a number that implies false precision. Understanding where your method’s detection limits fall is essential for knowing when your uncertainty budget is even applicable.

Monte Carlo Simulation as an Alternative to Classical Propagation

The classical law of propagation of uncertainty works well when the relationship between inputs and the final result is approximately linear and the uncertainties are relatively small. But real-world calculations in chemistry can involve nonlinear equations, asymmetric distributions, or situations where the simplifying assumptions behind the classical formula break down. Monte Carlo simulation offers a flexible alternative.

The idea is straightforward: instead of deriving a formula for how uncertainties combine, you let a computer do it by brute force. You specify the probability distribution for each input variable (normal, rectangular, triangular, or whatever matches your knowledge), then the computer randomly draws a value from each distribution, plugs them into the calculation, and records the output. Repeat this process a large number of times, often a million or more, and the spread of the output values directly shows you the uncertainty in the result. The approach has been applied to problems ranging from geochemical equilibrium calculations to marine carbon dioxide chemistry.

Comparisons between Monte Carlo and classical Gaussian propagation typically show excellent agreement when the assumptions behind the classical method are met. Software packages that implement both approaches have found that computed uncertainties agree within fractions of a percent for straightforward calculations, though Monte Carlo requires large sample sizes (on the order of 100,000 draws or more) to reliably match the Gaussian result within one percent. The real advantage of Monte Carlo emerges when the classical formula would be difficult to derive or when input distributions are not normal, making it a practical tool for complex analytical workflows.

Interlaboratory Variability and the Horwitz Function

Your own laboratory’s uncertainty budget captures the variability within your four walls, but results can look quite different when the same sample is measured by multiple laboratories using the same method. Interlaboratory variability is consistently larger than within-laboratory variability, because each lab brings its own subtle biases in equipment, reagent sourcing, analyst technique, and environmental conditions.

Food and environmental chemists have long recognized a striking empirical pattern known as the Horwitz function, which relates the reproducibility standard deviation from collaborative trials to the concentration of the analyte being measured. The relationship follows a predictable curve: as the concentration decreases, the relative standard deviation increases in a consistent way, regardless of the analyte, the matrix, or the method. This pattern has held remarkably well across decades of collaborative studies, and it provides a benchmark for judging whether a particular method’s between-lab variability is typical or unusually large.

The Horwitz function does not replace a proper uncertainty budget, but it offers a useful reality check. If your method’s reproducibility looks much worse than the Horwitz prediction, something may be wrong with the method or with how it is being implemented across sites. If it looks much better, you might want to scrutinize whether the collaborative trial was genuinely representative.

Practical Mistakes That Inflate or Hide Uncertainty

Knowing the formal framework is one thing; applying it well is another. Several common mistakes lead to uncertainty estimates that are either too small (giving false confidence) or unnecessarily large (wasting effort trying to improve a source of error that barely matters).

  • Ignoring correlated inputs: The propagation formula assumes that input uncertainties are independent. If two inputs are correlated (for example, if you used the same standard solution to calibrate two instruments), treating them as independent will underestimate the combined uncertainty. Correlation terms exist in the full propagation equation but are frequently left out.
  • Overlooking sample preparation: As discussed earlier, the steps between collecting a sample and presenting it to an instrument can dominate the uncertainty budget, yet many analysts estimate uncertainty only from the instrumental measurement onward.
  • Rounding too early: Rounding intermediate results before the final calculation can introduce rounding error that accumulates. Keep extra digits through the calculation and round only the final reported result.
  • Confusing precision with accuracy: A highly repeatable measurement (small random error) can still be badly wrong if a systematic bias is present. Reporting tight uncertainty intervals without having checked for bias, for example by analyzing a certified reference material, can be dangerously misleading.
  • Using too few replicates: A standard deviation calculated from three measurements is itself very uncertain. The fewer replicates you have, the less reliable your Type A uncertainty estimate is. When resources allow, more replicates give a more trustworthy picture of your method’s variability.

Software Tools and Automated Estimation

Calculating uncertainty by hand is feasible for simple cases, but as measurement models grow more complex, software becomes essential. The NIST Uncertainty Machine is a free web-based tool specifically designed for uncertainty propagation: you enter your measurement equation, specify the distribution and standard uncertainty of each input variable, and the tool calculates the combined uncertainty using both the classical GUM method and Monte Carlo simulation. It has been recommended as a teaching tool for chemistry students and is robust enough for many real analytical problems.

Domain-specific software also exists. In marine chemistry, for example, multiple software packages for computing carbonate system parameters now include built-in uncertainty propagation. Tests of these packages have shown that they agree with each other to better than 0.01 percent for calculated quantities and that their Monte Carlo results converge with Gaussian results when sample sizes are large enough. For clinical chemistry, guidelines such as ISO 20914 and the Nordtest approach provide structured procedures for estimating uncertainty from routine quality-control data, with the expanded uncertainty calculated by multiplying the combined uncertainty by a coverage factor of 2 for a 95 percent confidence interval.

The availability of these tools means there is less excuse than ever for skipping uncertainty estimation or for doing it poorly. Even if you are not deriving propagation formulas by hand, understanding what the software is doing under the hood helps you catch errors in your inputs and interpret the output critically.