Logistic regression output reports, for each variable in the model, how much that variable shifts the odds of a yes-or-no outcome occurring. The numbers you see are not probabilities or percentages but log-odds and their transformations, which is why the output looks cryptic at first. Once you learn to convert a few key values and know where to look for red flags, the whole table becomes readable in minutes.
What the Coefficient Table Is Telling You
The centerpiece of any logistic regression output is the coefficient table. Each row represents one predictor variable you put into the model. The column usually labeled “Estimate,” “Coef,” or “B” gives you the coefficient for that variable, which is the change in the log-odds of the outcome for every one-unit increase in the predictor. A positive coefficient means the variable is associated with higher odds of the outcome; a negative coefficient means lower odds.
Log-odds are not intuitive for most people, and that is perfectly fine. You rarely need to interpret them directly. Their main job is to serve as the raw material from which odds ratios are calculated. Think of the log-odds coefficient as the behind-the-scenes number and the odds ratio as the human-readable translation.
The table also includes a row for the intercept, sometimes labeled “(Intercept)” or “Constant.” The intercept represents the log-odds of the outcome when every predictor in the model equals zero. In many real-world models, zero on every predictor is a hypothetical scenario that does not correspond to an actual person or observation. If you are modeling whether a patient develops a complication, and the predictors include age and blood pressure, the intercept tells you the log-odds for a patient with age zero and blood pressure zero, which is meaningless in practice. So while the intercept is necessary for the math to work, you can usually ignore it when interpreting results. Where the intercept does matter is in making individual predictions: you need it to calculate the predicted probability for any specific combination of predictor values.
Turning Coefficients Into Odds Ratios
The single most useful thing you can do with a logistic regression coefficient is exponentiate it to get the odds ratio. Most software will do this for you automatically, but the principle is simple: raise the mathematical constant e (roughly 2.718) to the power of the coefficient. The result is the odds ratio for that predictor.
An odds ratio greater than 1 means the predictor is associated with increased odds of the outcome. An odds ratio less than 1 means it is associated with decreased odds. An odds ratio of exactly 1 means no association at all. So if a variable has an odds ratio of 3.0, the odds of the outcome are three times higher for every one-unit increase in that variable, holding everything else in the model constant. If the odds ratio is 0.5, the odds are cut in half.
The strength of logistic regression is that each odds ratio is adjusted for all the other variables in the model. If you include both smoking and age, the odds ratio for smoking reflects smoking’s association with the outcome after accounting for age, and vice versa. This is what makes logistic regression valuable for teasing apart which variables actually matter when several things are correlated with each other.1Biochemia Medica. Understanding logistic regression analysis
For categorical predictors with more than two levels (say, education coded as “high school,” “college,” and “graduate school”), the output picks one level as the reference category and reports odds ratios for the other levels compared to that reference. If “high school” is the reference, the odds ratio for “college” tells you how the odds differ for college-educated individuals compared to high-school-educated individuals. If you are not sure which level is the reference, look for the category that is missing from the table. That is the one the software chose as the baseline.
For continuous predictors, the odds ratio corresponds to a one-unit change. This means the practical meaning depends entirely on the scale of the variable. An odds ratio of 1.02 for age (measured in years) means a roughly 2% increase in odds per additional year of age, which adds up over decades. The same odds ratio of 1.02 for income measured in dollars would be trivially small. Always think about what “one unit” means for each predictor before deciding whether an odds ratio is large or small.
P-Values, Confidence Intervals, and Test Statistics
Next to each coefficient, you will find a p-value and usually a test statistic (often a Wald statistic or a z-value). The p-value tells you the probability of seeing a coefficient at least this extreme if the variable truly had no association with the outcome. The conventional threshold is 0.05: below it, the variable is considered statistically significant. Above it, you cannot confidently say it has an effect based on this data.
Confidence intervals for the odds ratio are more informative than p-values alone. A 95% confidence interval that does not cross 1.0 corresponds to a significant result. But the interval also tells you the plausible range of the true odds ratio. An odds ratio of 2.0 with a confidence interval of 1.1 to 3.6 is a very different finding from an odds ratio of 2.0 with a confidence interval of 1.8 to 2.2. Both are statistically significant, but the second is far more precisely estimated.
The most common test behind those p-values is the Wald test, which divides the coefficient by its standard error. Some software also reports a likelihood ratio test, which compares how well the model fits with and without each variable. These two approaches do not always agree, especially in small samples. Research comparing the Wald, likelihood ratio, and score tests has shown that they can have meaningfully different statistical power depending on the situation.2PubMed Central. Approximations of the power functions for Wald, likelihood ratio, and score tests and their applications to linear and logistic regressions In practice, when your sample size is large, the tests converge to similar answers. When it is small, the likelihood ratio test is generally considered the more reliable of the two.
Model Fit Statistics
Beyond individual coefficients, the output usually includes several numbers that describe how well the overall model fits the data. These do not tell you whether any single predictor matters; they tell you whether the model as a whole is doing a reasonable job of distinguishing between outcomes.
Two deviance values typically appear near the top or bottom of the output: null deviance and residual deviance. Null deviance measures how poorly a model with no predictors at all (just the intercept) fits the data. Residual deviance measures how poorly your model with all its predictors fits. If the residual deviance is substantially smaller than the null deviance, your predictors are collectively doing useful work. The difference between the two can be tested formally with a chi-square test.
The Akaike Information Criterion, or AIC, shows up in most software output and is useful for comparing models against each other. A lower AIC indicates a better balance between fit and complexity. It is not meaningful in isolation; it only matters when you compare two or more candidate models fit to the same data. If one model has an AIC of 340 and another has an AIC of 360, the first model is preferred.
You may also encounter pseudo-R² values, which attempt to summarize overall model performance in a way that resembles the R² from ordinary regression. Several versions exist (Nagelkerke, McFadden, Cox-Snell), and they do not have the same clean interpretation as the familiar R². A pseudo-R² of 0.3 does not mean the model explains 30% of the variance in the way you might be used to. It is better to think of pseudo-R² as a rough gauge: values closer to 0 suggest the predictors are not very helpful, and values closer to 1 suggest strong discrimination. Comparing one review of goodness-of-fit approaches, the likelihood ratio test, pseudo-R², and chi-square test produced similar conclusions about model fit when applied to the same data.3Asian Journal of Probability and Statistics. A Review of Some Goodness-of-Fit Tests for Logistic Regression Model
The Hosmer-Lemeshow Test and Its Limitations
One of the most commonly reported goodness-of-fit tests for logistic regression is the Hosmer-Lemeshow test. It works by dividing the observations into groups based on their predicted probabilities, then comparing how many events actually occurred in each group to how many the model predicted. A non-significant p-value (above 0.05) is taken to mean the model fits adequately. A significant p-value suggests the model’s predictions do not match the data well.
The test has a well-known weakness: it behaves differently depending on sample size. In very large datasets, even trivial departures from perfect fit will produce a significant result, leading you to reject a model that is practically useful. Research on this problem has led to proposed modifications that adjust for sample size so the test does not become overly sensitive as the dataset grows.4PubMed. Assessing the goodness of fit of logistic regression models in large samples: A modification of the Hosmer-Lemeshow test The opposite problem arises in complex models with many predictors: if the sample size is fixed, adding more predictors can cause the test to lose power and fail to detect genuine lack of fit.5PubMed Central. Improving the Hosmer-Lemeshow goodness-of-fit test in large models with replicated Bernoulli trials
The practical takeaway is that you should not rely on the Hosmer-Lemeshow test as your sole indicator of model quality. Use it alongside deviance, AIC, and measures of predictive discrimination like the area under the receiver operating characteristic curve (AUC). An AUC of 0.5 means the model is no better than a coin flip; 0.7 to 0.8 is considered acceptable discrimination in many fields; above 0.8 is strong.
Pitfalls That Can Distort Your Results
Even when the output looks clean, several common problems can make the numbers misleading. Knowing these pitfalls helps you spot suspicious output before drawing conclusions from it.
Multicollinearity
When two or more predictors in the model are highly correlated with each other, the standard errors of their coefficients get inflated. This makes the p-values unreliable and the confidence intervals unnecessarily wide. You might see a predictor that genuinely matters show up as non-significant simply because it is sharing its signal with another variable.6PubMed Central. Multicollinearity and misleading statistical results A variance inflation factor above 5 or 10 (the threshold varies by field) is a warning sign. The fix is usually to drop one of the correlated predictors or combine them.
Complete Separation
Sometimes a predictor perfectly separates the outcomes: every observation with a certain value of the predictor falls into one category. When this happens, the algorithm tries to push the coefficient toward infinity to perfectly classify those cases, and the output will show absurdly large coefficients with enormous standard errors. The model has not found an incredibly strong predictor; it has hit a wall. Firth’s penalized logistic regression is a well-established solution for this problem, as well as for the related issues of rare events and small sample sizes.7PubMed Central. Firth’s penalized logistic regression: A superior approach for analysis of data from India’s National Mental Health Survey, 2016 The penalized approach adds a small correction to the estimation procedure that pulls the coefficients back from extreme values.8Research Methods in Applied Linguistics. Dealing with complete separation and quasi-complete separation in logistic regression for linguistic data
Ignoring Nonlinearity
Standard logistic regression assumes that the relationship between each continuous predictor and the log-odds of the outcome is a straight line. If the true relationship is curved, say the risk rises sharply at first and then levels off, a straight-line assumption will misrepresent the effect and can produce misleading odds ratios. A systematic review of clinical prediction models found that roughly 85% of studies did not assess whether this linearity assumption held for their continuous predictors.9PubMed Central. Poor handling of continuous predictors in clinical prediction models using logistic regression: a systematic review Among the few that did check, the most common remedies were transforming the variable (for instance, using its logarithm) or fitting splines, which allow the relationship to bend at specified points. If you suspect a predictor has a non-linear relationship with the outcome, it is worth testing this before trusting the default output.
When Odds Ratios Overstate the Risk
One of the most consequential misinterpretations of logistic regression output happens when people treat odds ratios as if they were risk ratios. The two are similar only when the outcome is rare. As the outcome becomes more common, the odds ratio increasingly exaggerates the true change in risk. A classic analysis in JAMA demonstrated that when the outcome occurs in more than about 10% of the study population, the odds ratio can no longer serve as a reasonable stand-in for the relative risk. The more frequent the outcome, the larger the gap: an odds ratio above 1 will overestimate how much risk actually increases, and an odds ratio below 1 will overestimate how much it decreases.10PubMed. What’s the relative risk? A method of correcting the odds ratio in cohort studies of common outcomes
This matters most when you are reading output from a study where the outcome is not particularly rare, such as readmission rates, survey responses, or common conditions. If someone reports an odds ratio of 3.0 for a condition that affects 40% of the sample, the actual risk ratio could be considerably smaller. Several correction methods exist, including modified Poisson regression and the formula proposed in that JAMA paper. When you encounter logistic regression results applied to common outcomes, keep this discrepancy in mind before assuming the effect is as large as the odds ratio suggests.
Reading Output From Ordinal and Multinomial Models
Everything discussed so far applies to binary logistic regression, where the outcome has exactly two categories (yes/no, pass/fail, alive/dead). When your outcome has more than two categories, the output changes in important ways.
Ordinal logistic regression is used when the categories have a natural order, like “mild,” “moderate,” and “severe.” The output typically gives you one set of coefficients that applies across all the cutpoints between adjacent categories. This works because the model assumes that the effect of each predictor is the same regardless of where you draw the dividing line. That assumption, known as the proportional odds assumption, is not always met. When it is violated, the coefficients become harder to interpret and the model may fit poorly.11Mbeya University of Science and Technology Journal of Research and Development. Disparities in Methodology, Assumptions and Applications between Ordinal and Multinomial Logistic Regression: A Meta-Analysis Most software includes a test for this assumption, and if it fails, you may need to switch to a different approach.
Multinomial logistic regression handles outcomes with three or more categories that have no inherent order, like choosing among several transportation modes or diagnostic categories. The output is more complex: instead of one coefficient per predictor, you get one coefficient per predictor for each outcome category compared to a reference category. If you have four outcome categories and five predictors, the table will contain something like 15 coefficients (three non-reference categories times five predictors), each with its own odds ratio and p-value. The interpretation is the same as in binary regression, just repeated for each comparison. The tradeoff is that this produces a much larger table with many more numbers to make sense of, and statistical power drops because the data are spread across more comparisons.
Practical Tips for Scanning Output Quickly
When you sit down with a logistic regression table, whether from your own analysis or from a paper you are reading, a systematic approach saves time and prevents misreadings.
- Start with odds ratios, not coefficients: If the software provides both, go straight to the exponentiated values. They answer the question “how much do the odds change?” directly.
- Check confidence intervals before p-values: A narrow interval around a meaningful odds ratio tells you more than a p-value alone. A significant p-value paired with a huge confidence interval (say, 1.1 to 45.0) suggests the estimate is unstable.
- Look at the reference categories: For categorical predictors, the interpretation of the odds ratio depends entirely on what is being compared to what. Confirm the reference level before drawing conclusions.
- Note the sample size: Small samples can produce separation, inflated standard errors, and unreliable Wald tests. If the sample is small relative to the number of predictors (a common rule of thumb is at least 10 to 20 events per predictor), treat the results cautiously.
- Ask how common the outcome is: If more than about 10% of the sample experienced the outcome, the odds ratios may substantially overstate the risk ratios. This is especially relevant in cohort studies and surveys.
Getting comfortable with logistic regression output is less about memorizing formulas and more about building a mental checklist. The coefficient tells you the direction. The odds ratio tells you the magnitude. The confidence interval tells you the precision. The model fit statistics tell you whether the whole thing hangs together. And the pitfalls listed above tell you when not to trust what you see. With those pieces in place, even a dense output table becomes manageable.