AIC, or the Akaike Information Criterion, ranks competing models by balancing how well each fits the data against how many parameters it uses, and the model with the lowest AIC value in your set is considered the best-supported approximation. The number itself has no inherent meaning on a fixed scale; only the differences between AIC values across your candidate models tell you anything useful. That emphasis on differences rather than absolutes is where most of the practical interpretation lives, and it is also where most mistakes happen.
What the AIC Score Actually Measures
AIC estimates how much information is lost when a given model stands in for the process that actually generated your data. A model with a lower AIC loses less information, so it provides a better trade-off between accuracy and simplicity. Every additional parameter a model uses improves its fit to the data at least slightly, but AIC penalizes each extra parameter by adding 2 units to the score. If a parameter does not improve the fit enough to offset that 2-unit penalty, the model’s AIC goes up rather than down, and the parameter is not pulling its weight.
This penalty structure means AIC is not simply choosing the model that fits best. A heavily parameterized model will almost always fit the observed data more closely, but it risks capturing noise specific to your sample rather than the real underlying pattern. AIC tries to find the sweet spot: enough complexity to capture the signal, not so much that you are modeling randomness. The theoretical foundation traces back to information theory and the Kullback-Leibler divergence, a way of measuring how different two probability distributions are. In plain terms, AIC asks: “Of these candidate models, which one’s predictions will come closest to what reality would produce in new data?”
Reading Delta AIC Values
Because the raw AIC number is meaningless in isolation, you always work with differences. Subtract the lowest AIC in your set from each model’s AIC to get delta AIC (ΔAIC) for every candidate. The best model gets a ΔAIC of zero, and every other model’s ΔAIC tells you how much worse it is, in relative terms.
A widely used rule of thumb, popularized by Burnham and Anderson, groups models into rough tiers of support. Models within about 2 ΔAIC units of the best model have substantial support and cannot be confidently distinguished from the top model. Models with ΔAIC between roughly 4 and 7 have considerably less support. Models with ΔAIC greater than 10 have essentially no support relative to the best model. These guidelines are convenient but should not be treated as hard cutoffs. A model at ΔAIC 2.1 is not meaningfully worse than one at 1.9.
One way to think about how aggressive the ΔAIC < 2 threshold really is: when a more complex model differs from a simpler one by exactly one parameter, adding that parameter improves the AIC ranking as long as the parameter’s contribution to model fit exceeds a threshold equivalent to roughly p < 0.157 in a classical hypothesis test. That is far more permissive than the conventional p < 0.05 cutoff used in traditional significance testing.1PubMed Central. Practical advice on variable selection and reporting using Akaike information criterion This does not make AIC wrong; it reflects a different goal. AIC prioritizes predictive accuracy over strict parsimony, and that tradeoff intentionally leans toward including a parameter if there is a reasonable chance it improves prediction, even if you would not call it “statistically significant” by traditional standards.
The Uninformative Parameter Trap
The most common misinterpretation of AIC results involves models that land within 2 ΔAIC units of the best model not because they contain a genuinely useful predictor, but because they carry one extra parameter that does almost nothing. This is the uninformative parameter problem, and it trips up researchers constantly.
Here is how it works. Suppose your best model has three predictors. You also ran a model with those same three predictors plus a fourth. The fourth predictor reduces the model deviance by some tiny amount, not enough to overcome the 2-unit penalty AIC charges for the extra parameter. The result: the four-predictor model sits at ΔAIC of, say, 1.8. By the standard rule of thumb, that model has “substantial support.” But the reason it is close to the best model is not that the fourth predictor is doing meaningful work. It is close because the penalty for adding one useless parameter is at most 2 AIC units.2Journal of Wildlife Management. Uninformative parameters and model selection using akaike’s information criterion The parameter explains almost no variation and should not be interpreted as having a real effect.
The warning signs are straightforward. If a model within 2 ΔAIC of the best model differs from it only by adding one parameter, and that parameter’s confidence interval broadly overlaps zero or its effect size is negligible, you are likely looking at an uninformative parameter. The model is competitive on AIC purely by arithmetic, not because the extra variable matters. Researchers who interpret all models within ΔAIC < 2 as equally valid and then discuss every parameter in those models as if it has a meaningful effect end up drawing conclusions the data does not support.
Akaike Weights for Comparing Models
Delta AIC values give you a ranking and a rough sense of how far apart models are, but Akaike weights go a step further. They convert the ΔAIC values into a set of numbers between 0 and 1 that sum to 1, representing the relative likelihood of each model being the best approximation in your candidate set. A model with a weight of 0.72 is roughly 72% likely to be the best of the bunch, given the data and the set of models you compared.
Akaike weights make it easier to communicate results. Instead of saying “model A had the lowest AIC and model B was within 1.5 ΔAIC units,” you can say “model A had 60% of the weight and model B had 30%,” which immediately conveys how the support is distributed. When one model dominates with a weight above 0.90, the choice is clear. When several models share the weight more evenly, it signals genuine uncertainty about which structure best represents the data.
Where Akaike weights become tricky is model averaging. The idea sounds appealing: if you are not sure which model is best, average the predictions or parameter estimates across models, weighting each by its Akaike weight. For predictions, this often works well. For individual regression coefficients, though, averaging can produce nonsensical results when predictor variables are correlated with each other. The reason is that a coefficient’s meaning and scale can change depending on which other variables are in the model. Averaging a coefficient across models where it means slightly different things produces a number that does not have a clean interpretation.3Ecology. Model averaging and muddled multimodel inferences If you are model-averaging predictions for forecasting, you are on solid ground. If you are model-averaging individual coefficients to talk about the effect of a specific predictor, proceed with caution, especially when your predictors are correlated.
When to Use AICc Instead of AIC
Standard AIC assumes a large sample relative to the number of parameters. When that assumption breaks down, AIC tends to favor models that are more complex than the data can support, because the 2-unit-per-parameter penalty is not strong enough for small samples. The corrected version, usually written AICc, adds a stronger penalty that depends on sample size and the number of parameters. As the sample grows, the correction shrinks toward zero, and AICc converges to AIC.4Biometrika. Regression and time series model selection in small samples
A common rule of thumb is to use AICc when the ratio of sample size to parameters (n/K) is below about 40. In practice, many analysts just use AICc by default, since it does not hurt with large samples and helps with small ones. The interpretation of ΔAIC values and Akaike weights works the same way whether you are using AIC or AICc; only the scores themselves change.
The risk of ignoring the correction in small samples is real. With limited data, standard AIC can lead you to select a model with too many parameters, which means your parameter estimates will be noisy and your model may generalize poorly to new observations.
AIC Versus BIC
The Bayesian Information Criterion is AIC’s most common alternative, and the two frequently appear side by side in model-selection discussions. They look similar in structure but come from different philosophical starting points and sometimes disagree in practice, with AIC tending to favor more complex models and BIC favoring simpler ones.5Briefings in Bioinformatics. Sensitivity and specificity of information criteria
AIC tries to find the model that will predict new data most accurately. It does not assume the “true” model is among your candidates; it just wants the closest approximation. BIC, by contrast, tries to identify the true data-generating model, assuming one exists in your candidate set. Because BIC penalizes complexity more harshly (and the penalty increases with sample size), it tends to choose simpler models, especially with large datasets.
Which one you should use depends on what you are trying to do. If your primary goal is prediction and you are less concerned with finding the simplest possible explanation, AIC is the natural choice. If your priority is identifying a parsimonious, interpretable model and you believe the true model is relatively simple, BIC’s stronger penalty aligns with that goal.5Briefings in Bioinformatics. Sensitivity and specificity of information criteria In fields like ecology and biogeography, where the underlying processes are genuinely complex and no candidate model is likely to be the “true” one, AIC has become the more popular choice. In areas where researchers believe true effects are sparse and simple models dominate, BIC gets more use.
One practical nuance: when there is substantial unobserved variability in the data that your models do not capture, AIC and AICc can struggle. Under those conditions, BIC’s heavier penalty often performs better because it is less likely to select a complex model that is fitting noise driven by that hidden variability.6Methods in Ecology and Evolution. The relative performance of AIC, AIC C and BIC in the presence of unobserved heterogeneity If you suspect your candidate models are all missing an important source of variation, keep this in mind.
AIC Only Compares the Models You Give It
A point that sounds obvious but causes real problems: AIC ranks the models in your candidate set. It does not tell you that the best model in your set is actually a good model in any absolute sense. If all your candidates are terrible approximations of the data-generating process, AIC will still point to the least terrible one and give it a ΔAIC of zero. You could have a “winning” model that predicts the data poorly.
This means AIC results need to be paired with checks on absolute model fit. Examining residual plots, checking whether predictions are reasonable, and looking at measures of overall goodness of fit (like R-squared for regression models, or deviance-based diagnostics for other model types) are still essential. AIC handles the question “which of these models is best?” but not “is the best one good enough?”
It also means the candidate set itself matters enormously. If you fail to include a model structure that captures an important feature of the data, AIC cannot rescue you. The quality of your conclusions is bounded by the quality of your model set. Thoughtful model construction based on domain knowledge before you ever compute AIC is more important than any post hoc ranking exercise.
Handling Overdispersed or Structured Data
Standard AIC assumes that the data points are independent and that the model’s error distribution is correctly specified. When those assumptions fail, AIC can behave unpredictably. One common violation is overdispersion, where the data show more variability than the assumed distribution predicts. In count data modeled with a Poisson distribution, for instance, overdispersion is widespread. Under overdispersion, AIC tends to select models that are more complex than necessary, because the extra parameters soak up variance that actually comes from the distribution being mis-specified rather than from a missing predictor.7Methods in Ecology and Evolution. Model selection with overdispersed distance sampling data
The standard fix is QAIC, which adjusts the AIC calculation by an estimated overdispersion factor. Think of it as recalibrating the penalty to account for the fact that your data are more variable than the model expects. If you are working with count data or binary data and suspect overdispersion, checking for it and switching to QAIC is a straightforward safeguard against selecting needlessly complex models.
A different structural issue arises with mixed models that include random effects. In those settings, you can compute AIC based on either the marginal likelihood (which integrates over the random effects) or the conditional likelihood (which conditions on estimated random-effect values). These two versions of AIC can disagree about which model is best. The marginal version tends to favor simpler models that drop random effects, while the conditional version can be biased by uncertainty in estimating the random-effects structure.8Biometrika. On the behaviour of marginal and conditional AIC in linear mixed models If you are comparing mixed models, being explicit about which likelihood you are using for AIC, and understanding that the choice affects the outcome, is important.
Reporting AIC Results Clearly
How you report AIC results matters almost as much as how you compute them. A bare statement like “model 3 was the best model (AIC = 234.7)” tells the reader very little. Good reporting involves presenting the full candidate set, or at minimum all models within about 8 ΔAIC units of the best, along with the number of parameters, the log-likelihood, the ΔAIC value, and ideally the Akaike weight for each model.9PLOS ONE. On the prevalence of uninformative parameters in statistical models applying model selection in applied ecology This transparency lets readers spot uninformative parameters and evaluate whether the model ranking makes substantive sense.
Parameter estimates with confidence intervals for each model in the competitive set should also appear, either in the main text or a supplement. Without them, it is impossible for a reader to judge whether a parameter in a “competitive” model (ΔAIC < 2) is genuinely informative or just coming along for the ride. Graphical presentations of modeled relationships can help too, especially for making it visually apparent when two models produce nearly identical predictions despite differing in structure.
A survey of ecological studies found that uninformative parameters show up frequently in published model-selection results, often because authors treat all models within the ΔAIC < 2 window as equally valid and proceed to discuss every parameter those models contain.9PLOS ONE. On the prevalence of uninformative parameters in statistical models applying model selection in applied ecology This is not a minor formatting issue; it leads to real misinterpretation of which variables matter. The fix is straightforward: when a model sits within 2 ΔAIC of the best and differs only by one added parameter, check that parameter’s estimate and confidence interval before claiming it has a meaningful effect.
AIC in a Broader Model-Selection Landscape
AIC occupies one corner of a much larger model-selection landscape. Cross-validation, where you repeatedly fit a model on part of the data and test it on the rest, directly estimates predictive performance without relying on the assumptions that underpin AIC. For large, complex models where AIC’s assumptions about the penalty term may not hold well, cross-validation can be a more robust (if more computationally expensive) alternative.
In Bayesian analysis, criteria like the Deviance Information Criterion (DIC) and the Widely Applicable Information Criterion (WAIC) serve analogous roles. WAIC, in particular, is sometimes called the “Bayesian AIC” because it estimates predictive accuracy in a way that accounts for the full posterior distribution of the parameters rather than relying on a point estimate.10PubMed. Comparing DIC and WAIC for multilevel models with missing data If you are working in a Bayesian framework, these tools are more natural fits than AIC, which was designed for frequentist maximum-likelihood estimation.
The rise of information-theoretic model selection as a mainstream approach, particularly in ecology, represented a genuine shift away from null-hypothesis significance testing as the default.11Journal of Applied Ecology. Information theory and hypothesis testing: a call for pluralism But AIC is not a universal replacement for hypothesis tests. It answers a different question: not “is this effect statistically distinguishable from zero?” but “does including this effect make for a better predictive model?” Those two questions can give different answers about the same variable. A variable with a genuine but small effect might fail a significance test at p < 0.05 but still improve AIC. Conversely, a variable could be “significant” by a p-value criterion but not improve AIC when a simpler model already captures most of the pattern. Understanding which question you are asking is the most important step in choosing the right tool.
Spatial Autocorrelation and Other Structural Violations
Switching from hypothesis testing to AIC-based model selection does not solve all the statistical problems that plague a dataset. Spatial autocorrelation is a good example. When observations that are close together in space tend to be more similar to each other than distant ones, the effective sample size is smaller than the nominal sample size. This means the data contain less independent information than AIC “thinks” they do. The result is that AIC can favor overly complex models for the same reason it does with overdispersed data: extra parameters appear to improve fit, but they are partly just soaking up structure that belongs to the spatial pattern rather than to the predictors.12Global Ecology and Biogeography. Model selection and information theory in geographical ecology
This does not mean AIC is useless for spatially structured data. It means you should not expect the shift from p-values to AIC to magically fix spatial correlation issues. The standard advice is to account for spatial structure in the model itself, through spatial random effects or explicit correlation structures, and then use AIC to compare among models that all address the spatial issue. Using AIC to compare models that ignore spatial autocorrelation is asking the criterion to do a job it was not designed for.
The same logic extends to other forms of non-independence: repeated measures on the same subject, hierarchical or clustered data, and time-series observations. When the independence assumption is violated, AIC’s penalty term may be miscalibrated. The fix is almost always the same: build the dependence structure into the model before comparing models with AIC, rather than hoping AIC will sort it out on its own.