Nested design refers to a study structure in which smaller units of observation sit inside larger ones, like Russian dolls: cells measured within animals, animals housed within cages, cages placed within rooms. Recognizing and properly handling this layering is one of the most consequential decisions in biological research, because ignoring it can inflate false-positive rates from the intended five percent to as high as eighty percent. The concept applies everywhere from rodent toxicology and single-cell sequencing to river ecology and multi-site clinical trials, yet a surprising share of published studies still get it wrong.
Why Ignoring Nesting Inflates False Positives
Most conventional statistical tests assume that every observation in your dataset is independent of every other. When data are nested, that assumption breaks down. Cells taken from the same animal tend to look more alike than cells taken from different animals. Animals raised in the same cage share environmental exposures their cage-mates across the room do not. This within-group similarity means that treating every individual measurement as a fully independent data point overstates how much information you actually have. You think your sample size is large, but the effective sample size is much smaller.
The practical fallout is severe. A study published in Nature Neuroscience showed that ignoring nested dependency can push the probability of a false positive far above the nominal five-percent threshold, reaching as high as roughly eighty percent in some scenarios.1Nature Neuroscience. A solution to dependency: using multilevel analysis to accommodate nested data Separate simulation work confirmed that cluster-related variation, when left unmodeled, produces false-positive rates in the range of twenty to fifty percent.2PubMed Central. Multilevel analysis quantifies variation in the experimental effect while optimizing power and preventing false positives In other words, a researcher might celebrate a “significant” treatment effect that is really just the natural similarity among observations sharing a common cluster.
The Litter Effect in Animal Research
One of the clearest and most well-documented examples of nested design problems comes from rodent studies. When a pregnant dam is exposed to a drug, toxin, or stressor, the entire litter shares that exposure. The pups are not independent replicates of the treatment; they are repeated measurements within the same experimental unit, which is the dam. This is the litter effect, and it bites hard.
Littermates resemble each other across morphological, biochemical, and behavioral measures more than they resemble pups from other litters, even when no treatment has been applied.3PubMed Central. Controlling litter effects to enhance rigor and reproducibility with rodent models of neurodevelopmental disorders When researchers treat each pup as an independent data point instead of accounting for litter membership, they dramatically overcount their sample size. A study examining research on valproic acid, a common model for autism-like behaviors in rodents, found that litter effects accounted for up to sixty-one percent of the variation in behavioral outcomes, which was larger than the treatment effects themselves. Only about nine percent of the studies reviewed correctly identified the litter as the experimental unit.4PubMed Central. Improving basic and translational science by accounting for litter-to-litter variation in animal models That means the vast majority of those studies were making statistical inferences based on a unit of analysis that was too granular, inflating apparent significance.
More recent work on acid-sensing ion channel knockout mice reinforced the same lesson: analyzing individual pups as independent data points rather than accounting for their litter of origin leads to Type I errors and should be avoided.5PubMed Central. Considering Litter Effects in Preclinical Research: Evidence from E17.5 Acid-Sensing Ion Channel 2a Knockout Mice Exposed to Acute Seizures The fix is conceptually simple: either average outcomes within each litter and use the litter mean as your data point, or include litter as a random effect in your statistical model. Either way, the litter, not the pup, is the unit that gets counted toward your sample size for the treatment comparison.
How Common Is the Mistake
The litter-effect problem is not an isolated oversight. Stuart Hurlbert coined the term “pseudoreplication” in 1984 to describe the practice of treating non-independent observations as independent replicates, and his survey of 176 ecological field experiments found that pseudoreplication occurred in about twenty-seven percent of them. Among the subset of studies that used inferential statistics, the rate was forty-eight percent.6Ecological Monographs. Pseudoreplication and the Design of Ecological Field Experiments Marine benthos and small-mammal studies were especially prone to the error. That paper has been cited thousands of times and remains one of the most-read methodological critiques in ecology, yet the problem persists across biological disciplines.
More recently, researchers have argued that much of the confusion stems from a failure to distinguish among three types of units: biological units (the organisms or cell lines of interest), experimental units (the smallest physical entity that can be independently assigned to a treatment), and observational units (whatever is actually measured).7PubMed Central. What exactly is ‘N’ in cell culture and animal experiments? When you measure three fields on a histological slide from one mouse, you have three observational units but still just one experimental unit. Genuine replication requires more mice, not more microscope fields, unless the question you are asking is specifically about within-mouse variability.
Nested Versus Crossed Designs
Not all hierarchical data structures are the same. In a nested design, each lower-level unit belongs to exactly one higher-level group: pups belong to one litter, and that litter belongs to one treatment. In a crossed design, lower-level units can appear in combination with every level of the other factor: the same observer might score animals in every treatment group, or the same genotype might be tested in every environment.
The critical difference is what happens to the interaction between factors. In a crossed design, you can estimate the interaction directly because each combination exists. In a nested design, the interaction is inseparable from the main effect of the nested factor; the two get lumped together. Data can be nested either because the study was set up that way on purpose, or because it would have been impractical or meaningless to cross the factors. When nesting is a deliberate design choice even though crossing was feasible, the pooling of variances needs to be clearly acknowledged, because the researcher has chosen to give up the ability to estimate that interaction.8Methods in Ecology and Evolution. Nested by design: model fitting and interpretation in a mixed model era When nesting is a natural feature of the system, such as individual fish living in particular river sections, the pooling is less problematic because the interaction simply does not exist in a meaningful sense.
Understanding whether your data are nested or crossed determines how you build your statistical model. Code them incorrectly and you either fail to estimate variation you could have estimated, or you pretend to estimate variation that does not exist in your data. Neither error is harmless.
Beyond the Lab Bench
Nested structures show up well beyond cages and litters. In field ecology, organisms are sampled within habitats, habitats within sites, sites within regions. A study of fish, mussels, and macroinvertebrates in a Central European river system used a nested sampling design spanning five rivers, with four river sections per river and five sampling sites per section, each split into pool and riffle habitats.9Journal of Biogeography. Using a Nested Sampling Design Across Spatial Scales to Gain Insights Into Distribution Patterns of Fishes, Mussels and Macroinvertebrates in a Riverine System That deliberate layering let researchers partition how much of the variation in species distribution was driven by broad river identity versus local habitat type. Without the nested structure, those spatial scales would have been confounded.
In clinical research, the same logic applies to multi-site trials. The STRIDE trial, which tested a multifactorial fall-prevention program, enrolled patients nested within clinics, which were themselves nested within larger health care systems. Randomization happened at the clinic level, making the study what methodologists call a subcluster randomized trial.10PubMed Central. Designing three-level cluster randomized trials to assess treatment effect heterogeneity Patients within the same clinic share the same clinical culture, staffing, and patient-education norms. Treating every patient as independent would overstate the precision of the treatment-effect estimate, just as treating every pup as independent overstates precision in a rodent study.
Nesting in Single-Cell and Genomics Experiments
High-throughput molecular biology adds another layer of complexity. In a single-cell RNA sequencing experiment, thousands of cells might be captured and sequenced on one plate, and a different batch of cells on another plate. Cells from the same plate share technical artifacts: capture efficiency, amplification bias, and sequencing depth. If you run one plate per biological condition, it becomes impossible to tell whether differences between conditions reflect genuine biology or plate-to-plate technical variation.
Researchers studying this problem designed their experiments with multiple technical replicates per biological individual, which allowed them to directly estimate the batch effect associated with independent preparations.11Scientific Reports. Batch effects and the effective design of single-cell gene expression studies The lesson parallels the animal research story: if your biological conditions are completely confounded with your technical batches, no amount of statistical correction after the fact can cleanly separate the two. The design itself has to include replication at the right level.
Using Mixed Models to Handle the Hierarchy
The dominant tool for analyzing nested biological data is the linear mixed model. In a mixed model, fixed effects represent the variables whose specific levels you care about, typically your treatment groups. Random effects represent grouping factors whose specific levels are drawn from a larger population, things like individual animals, litters, cages, clinics, or sequencing plates. By including random effects, the model acknowledges that observations within a cluster are correlated and adjusts its estimates of uncertainty accordingly.
Mixed models have become increasingly common in ecology and biology more broadly.12PubMed Central. A brief introduction to mixed effects modelling and multi-model inference in ecology They can also be extended to partition variation across multiple hierarchical levels simultaneously. In ecoimmunology, for instance, multi-response mixed models have been used to separate immune-response variation into within-individual, among-individual, and among-species components.13Integrative and Comparative Biology. Testing Hypotheses in Ecoimmunology Using Mixed Models: Disentangling Hierarchical Correlations Getting the random-effects structure right matters enormously: specify too few levels and you get inflated false positives; specify the wrong structure and your variance estimates become unreliable.
One practical tip from guidelines on laboratory animal experiments: when multiple observations come from a single experimental unit, a nested analysis of variance can estimate the “components of variance” at each level. If the biggest chunk of variation lives at the level of microscope fields within an animal, you gain more statistical power by examining more fields per animal than by adding more animals, assuming both options cost about the same.14ILAR Journal. Guidelines for the Design and Statistical Analysis of Experiments Using Laboratory Animals This is a concrete example of how understanding the nesting structure helps you spend your resources wisely rather than just defaulting to “use more animals.”
Where to Invest Your Sample Size
One of the most practical decisions a researcher faces with nested designs is how to allocate effort across levels. Should you sample more sites, more plots within sites, or more quadrats within plots? More patients, more clinics, or more health systems? The answer depends on where most of the variation lives and what each additional unit costs.
In environmental monitoring, a study of larval surveys found that increasing the sample size at the lowest level of the hierarchy was the most realistic way to boost statistical power, and that fixed-plot designs had greater power than random-plot designs.15PubMed. Using variance components to estimate power in a hierarchically nested sampling design That result is not universal, though. When most variation sits at the cluster level, adding more clusters (the top of the hierarchy) will typically yield more power than packing more observations into existing clusters. The key is to estimate variance components from pilot data or previous studies, then use those estimates to plan your allocation.
For clinical trials with nested structures, the optimal allocation ratio depends on the research question. Balanced designs, where you assign equal numbers to each treatment arm, tend to be optimal or highly efficient for most comparisons. The main exception is when you want to contrast one specific treatment arm against all others, in which case an unbalanced allocation performs better.16PubMed. Efficient treatment allocation in two-way nested designs
Reporting Transparency
Even when researchers handle nesting correctly in their analysis, the methods sections of published papers often fail to communicate what was done. The ARRIVE guidelines 2.0, which set reporting standards for animal research, explicitly address this gap. They note that hierarchies such as cells within animals and mitochondria within cells, or cages within rooms and animals within cages, make determining the true sample size difficult and require either including the clustering factor in the statistical model or aggregating outcomes to the appropriate level.17PLoS Biology. Reporting animal research: Explanation and elaboration for the ARRIVE guidelines 2.0
Transparent reporting means stating clearly what the experimental unit was, how many experimental units were used, and whether nesting or clustering was accounted for in the analysis. Without that information, reviewers and readers cannot judge whether the reported p-values and confidence intervals are trustworthy. The nine-percent compliance rate in the valproic acid literature mentioned earlier suggests that the field has a long way to go on this front.
Bayesian Approaches to Hierarchical Data
Frequentist mixed models are the workhorse, but Bayesian hierarchical models offer an alternative framework that some researchers find more intuitive for deeply nested data. In a Bayesian setup, each level of the hierarchy gets its own probability distribution, and information flows between levels: estimates for a specific cluster are “shrunk” toward the overall mean when data within that cluster are sparse. This partial pooling can be particularly useful when some groups have very few observations.
Bayesian hierarchical methods have been applied across a range of biological problems. In malaria research, an R software package called “bhrcr” uses Bayesian hierarchical regression to estimate parasite clearance rates across patients while accounting for covariates and the phases of parasite decline.18PubMed Central. Malaria parasite clearance rate regression: an R software package for a Bayesian hierarchical regression model In population genetics, a fast Bayesian hierarchical clustering tool called fastbaps can handle alignments of over 110,000 sequences, making it feasible to apply model-based clustering at scales that older methods could not reach.19Nucleic Acids Research. Fast hierarchical Bayesian analysis of population structure These are specialized applications, but they illustrate how the same core principle of modeling variation at multiple levels extends well beyond the traditional nested analysis of variance.
Common Misconceptions About Nested Designs
A few misunderstandings come up repeatedly. The first is that adding more observations per cluster somehow compensates for having few clusters. It does not. If you have only three litters per treatment and measure twenty pups per litter, you still have a sample size of three for the treatment comparison, no matter how sophisticated your model is. The within-cluster observations help you estimate within-cluster variation more precisely, but they do not give you additional information about between-cluster differences.
A second misconception is that nesting is always obvious from the data structure. Sometimes it is not. If you rotate lab technicians across treatment groups so that each technician processes samples from every group, technician is a crossed factor. If each technician handles only one group, technician is nested within treatment. The same factor can be nested or crossed depending on how the experiment was run, and the distinction changes which model you should fit.8Methods in Ecology and Evolution. Nested by design: model fitting and interpretation in a mixed model era
A third is the belief that post-hoc statistical corrections can rescue a confounded design. If your biological conditions are completely nested within batches, with no replication of conditions across batches, then no correction can separate the treatment effect from the batch effect. The confound is structural, baked into the design itself, and no amount of modeling after data collection can undo it. This is especially relevant for high-throughput experiments where running an extra batch feels expensive but is sometimes the only way to make the data interpretable.
When Nesting Is a Feature, Not a Bug
It is worth noting that hierarchical structure in data is not inherently a problem. In many cases, the nested structure is exactly what makes the science interesting. An ecologist studying how species composition changes across spatial scales needs the nested design to partition variation between local habitat, reach, and river. An immunologist studying how immune responses vary within and among individuals needs the hierarchy to ask whether the interesting variation is happening at the cellular level or the organism level.13Integrative and Comparative Biology. Testing Hypotheses in Ecoimmunology Using Mixed Models: Disentangling Hierarchical Correlations A clinical trialist designing a multi-site study can exploit the nesting to ask whether treatment effects vary across sites, which is valuable information for deciding whether an intervention will generalize.
The trouble only arises when the nesting is present but unacknowledged, leading to pseudoreplication and inflated certainty. When it is acknowledged and modeled appropriately, nested data give you richer answers than flat data ever could. You learn not just whether a treatment works on average, but how variable its effect is across the groups in your hierarchy, and where in the hierarchy the interesting action is happening.