Can the Akaike Information Criterion (AIC) Be Negative?

AIC can absolutely be negative, and there is nothing wrong with your model or your software when it happens. A negative AIC is a perfectly normal arithmetic result of the way the criterion is calculated. The value itself, whether positive or negative, carries no special meaning on its own. What matters is how one model’s AIC compares to another’s, because AIC is a tool for ranking competing models, not for grading any single model in isolation.

How the Formula Produces a Negative Number

AIC is calculated as two times the number of parameters in the model, minus two times the log-likelihood of the data given that model. The first part (twice the parameter count) is always a positive number and acts as a penalty for complexity. The second part (twice the log-likelihood) captures how well the model fits the data. When the log-likelihood is a large positive number, subtracting twice its value can easily drag the whole expression below zero.

A quick illustration makes this concrete. Imagine a model with 7 parameters and a log-likelihood of 70. The penalty term is 14, and the fit term is 140, so AIC equals 14 minus 140, which is −126. Nothing exotic happened. The log-likelihood was simply large enough relative to the number of parameters that the subtraction produced a negative result. This is routine in many types of statistical modeling.

Why AIC Values Have No Standalone Meaning

One of the most common sources of confusion is treating AIC like a score where “closer to zero” or “positive” means something good. It does not. The raw number on its own tells you almost nothing useful about model quality. AIC was designed to estimate the relative amount of information lost by each candidate model, so it only becomes informative when you compare two or more models fitted to the same dataset. A model with an AIC of −200 is not inherently better than one with an AIC of 50; those numbers only matter if both models were fitted to the same data and you are comparing them head to head. The model with the lower AIC is preferred because it represents a better trade-off between fit and complexity.

Research on AIC’s theoretical foundations reinforces this point. The precise value of AIC has no direct interpretation; it is the difference between AIC values across candidate models that guides selection.1arXiv. Estimating a difference between Kullback-Leibler risks by a normalized difference of AIC A single AIC number cannot tell you whether a model is “good enough” in any absolute sense. It can only tell you which model, among those you tested, wastes the least information.

When Negative AIC Values Tend to Show Up

Negative AIC values appear most often when the likelihood of the data under the model is high, which makes the log-likelihood a large positive number. This is common in several situations:

  • Continuous distributions: When your model uses a probability density function rather than a probability mass function, the density values at observed data points can exceed 1. That means the log-likelihood can easily be positive and large, pushing AIC into negative territory. Many regression models with continuous response variables fall into this category.
  • Well-fitting models with few parameters: A parsimonious model that happens to capture the data-generating process well will have a high log-likelihood and a small penalty term. The gap between the two terms widens, and the result is a strongly negative AIC.
  • Large sample sizes: More data points generally push the total log-likelihood further from zero, often in the positive direction for well-specified models. With enough observations and a reasonable model, negative AIC is the norm rather than the exception.

Conversely, AIC tends to stay positive when the log-likelihood is small or negative, which happens with certain discrete models, poorly fitting models, or models with many parameters. But again, the sign tells you nothing diagnostic. A positive AIC of 300 and a negative AIC of −50 could both be perfectly fine models for their respective datasets.

The Fit-Complexity Trade-Off in Plain Terms

AIC achieves parsimony by balancing two competing goals. One component measures model fit through the deviance, and the other penalizes model complexity through the parameter count.2PubMed Central. Practical advice on variable selection and reporting using Akaike information criterion A model that fits the data brilliantly but uses a large number of parameters gets a heavier penalty. A model with few parameters that fits poorly gets a large deviance term. The AIC is the balance point between these two forces, and the model with the lowest AIC is considered the best compromise.

The penalty term grows linearly with the number of parameters. Adding one parameter to a model increases AIC by 2 unless the log-likelihood improves by more than 1. That is a fairly lenient penalty, which is why AIC sometimes favors slightly more complex models compared to other criteria. The key practical consequence: when your candidate models all produce negative AICs, the most negative one is still the preferred model, because “lowest AIC wins” applies regardless of sign.

Delta AIC and Model Ranking

Because the raw AIC number is not interpretable on its own, practitioners typically work with differences between models. You take the AIC of each candidate model and subtract the AIC of the best model (the one with the lowest value). The best model gets a difference of zero, and every other model gets a positive number reflecting how much worse it is.

These differences are the basis for practical decision-making. A difference of less than about 2 is generally taken to mean two models are essentially equivalent, and you cannot distinguish between them on the basis of AIC alone. A difference of 4 to 7 suggests the worse model has noticeably less support. A difference greater than 10 suggests the worse model is very unlikely to be the best approximation. None of this reasoning changes when the raw AIC values happen to be negative. If your best model has an AIC of −340 and the next-best has −332, the difference is 8, and that interpretation is exactly the same as if the values were 340 and 348.

This is worth emphasizing because people who see a negative AIC for the first time sometimes start hunting for bugs in their code. They re-check data entry, toggle software settings, or try different optimization algorithms, all because the number “looks wrong.” It is not wrong. The sign is irrelevant to the comparison.

Akaike Weights for Easier Interpretation

If raw differences between negative numbers still feel unintuitive, Akaike weights offer an alternative way to present the results. Akaike weights convert the set of AIC differences into proportions that sum to 1, so each candidate model gets a number between 0 and 1 representing its relative probability of being the best model in the set. A model with a weight of 0.85 has an 85% chance of being the best approximation among the candidates you tested.

Weights sidestep the confusion around negative values entirely, since the output is always a proportion. They also make it easier to communicate results to audiences who are not familiar with information criteria. If you are writing a report for non-statisticians and you suspect the negative AIC values will raise eyebrows, presenting Akaike weights instead can save you a round of explaining that negative numbers are fine.

Can You Compare AIC Values Across Different Datasets?

No, and this is one of the most common mistakes in practice. AIC values are only comparable when the models were fitted to exactly the same dataset. If you change the sample size, transform the response variable, drop observations with missing data, or switch from one response variable to another, the AIC values from the two analyses live on different scales and cannot be compared.

This restriction matters more than it sounds. In practice, different software packages sometimes handle missing data differently by default, dropping different rows for different models. If one model was effectively fitted to 980 observations and another to 1,000 because of how missing values were handled, their AIC values are not comparable even though both analyses started from the same spreadsheet. This is a pitfall that catches experienced analysts, not just beginners.

The same logic extends to models with different likelihood functions. If you fit a linear regression and a logistic regression to the same data (treating the response as continuous in one case and binary in the other), the two AIC values are built on fundamentally different likelihoods and cannot be ranked against each other. AIC is a tool for comparing models that share the same data and the same response structure.

The Small-Sample Correction and Negative Values

When the number of observations is small relative to the number of parameters, the standard AIC tends to favor overly complex models. A corrected version, usually called AICc, adds an extra penalty term that grows as the ratio of parameters to sample size increases. AICc converges to the standard AIC as the sample size grows large, so for big datasets the difference is negligible.

AICc can also be negative, for the same reason the standard version can. The correction adds a positive quantity to the standard AIC, so AICc is always at least as large as AIC for the same model. If the standard AIC is deeply negative, the corrected version will be less negative but can still easily remain below zero. All the same comparison rules apply: rank models by AICc, prefer the lowest value, and ignore the sign.

A common rule of thumb is to use AICc whenever the sample size divided by the number of parameters is less than about 40. Since the correction never hurts (it just converges to zero when unnecessary), some analysts use AICc by default in all situations, which is a defensible choice.

How BIC Differs and Whether It Can Also Be Negative

The Bayesian Information Criterion (BIC) is the other widely used model selection criterion. It has the same basic structure as AIC but imposes a heavier penalty for each parameter, because the penalty scales with the natural logarithm of the sample size rather than being a fixed value of 2. For any dataset with more than about 8 observations, BIC penalizes extra parameters more severely than AIC does.

BIC can also be negative, and for the same fundamental reason: if the log-likelihood is large enough to overwhelm the penalty term, the result is negative. In practice, BIC tends to be less negative than AIC for the same model because the penalty is bigger, but the sign carries no special meaning for BIC either. As with AIC, only the relative ranking of BIC values across candidate models matters.

The two criteria sometimes disagree on which model is best. AIC tends to pick slightly more complex models, while BIC tends to pick simpler ones. Neither is universally superior. If your goal is prediction, AIC is often the better guide. If your goal is identifying the “true” model (assuming one exists among your candidates), BIC has theoretical advantages. When both criteria agree, you can be more confident in the selection. When they disagree, the models involved are usually close enough that the practical difference is small.

Software Quirks That Create Confusion

Different software packages sometimes report AIC values that do not match even when fitted to the same data with the same model. This happens because some packages include constant terms in the log-likelihood calculation that others omit. These constants do not affect model ranking within a single software environment, since the same constant is added or subtracted from every model, but they can shift the absolute AIC value by a large amount.

For example, you might fit the same linear regression in two different programs and get an AIC of −126 in one and an AIC of 340 in the other. Neither is wrong. One is including certain normalizing constants in the likelihood, and the other is dropping them because they cancel out during model comparison anyway. The lesson: never compare raw AIC values produced by different software packages unless you have verified that they compute the likelihood in the same way.

Some software also reports AIC per observation (dividing by the sample size), which produces smaller numbers and can further confuse cross-software comparisons. If your AIC values seem suspiciously small, check whether the software is normalizing by sample size.

Common Misconceptions About AIC

Several persistent misunderstandings circulate about AIC, especially among researchers who use it as a routine tool without digging into the underlying logic.

  • Lower AIC means a good model: No. A model can have the lowest AIC among your candidates and still be a terrible model. If all your candidate models are badly specified, AIC will faithfully pick the least bad one, but “least bad” is not the same as “good.” AIC cannot tell you whether any of your models adequately captures the data-generating process.
  • AIC should be close to zero: There is no target value. An AIC of −10,000 is not worse or better than −10 in any absolute sense. Only comparisons between models carry meaning.
  • Negative AIC means overfitting: The sign of AIC has no relationship to overfitting. Overfitting is penalized by the parameter-count term, and that penalty operates the same way whether the final number is positive or negative.
  • You can compare AIC to a chi-squared distribution: AIC is not a test statistic and has no associated p-value. It does not follow a chi-squared or any other standard distribution. Trying to assess “significance” of an AIC value is a category error.

These misconceptions often stem from conflating AIC with goodness-of-fit statistics that do have interpretable absolute values, like R-squared. AIC serves a fundamentally different purpose: it estimates relative information loss, not absolute model quality.2PubMed Central. Practical advice on variable selection and reporting using Akaike information criterion

Reporting AIC in a Paper or Report

If you are writing up results that involve AIC-based model selection, a few practices make the presentation clearer for readers. Report the AIC for every candidate model, not just the winner. Present the delta-AIC values alongside the raw values, since these are what readers need to assess the strength of evidence. If you have many candidate models, a table sorted by AIC with a column for delta-AIC and a column for Akaike weights is the most compact way to convey the results.

When the raw AIC values are negative, there is no need to explain or defend this. If you anticipate a reviewer or reader who might question it, a brief parenthetical noting that AIC values are relative and their sign is irrelevant to model comparison is sufficient. The far more important thing to report is the set of candidate models and why you chose them, since AIC can only choose among the models you give it. A thoughtfully constructed candidate set matters more than any single AIC value.

It is also good practice to report how many observations were used in each model fit, to reassure readers that the AIC values are indeed comparable. If any models were fitted to slightly different subsets of the data due to missing values, flag this explicitly, because the AIC values from those models cannot be ranked against the rest.1arXiv. Estimating a difference between Kullback-Leibler risks by a normalized difference of AIC