Building a statistical model follows a sequence that starts well before any numbers get crunched and continues well after you get your first results. The process moves from defining a precise question, through data exploration and cleaning, into choosing and fitting a model, and then into validation and ongoing monitoring. Each step shapes what comes next, and skipping one almost always creates problems downstream. The specifics change depending on whether you are trying to explain why something happens or predict what will happen next, but the skeleton of the process stays remarkably consistent.
Define the Question First, Not the Model
The most common mistake in statistical modeling is starting with a technique rather than a question. Researchers and analysts frequently choose a method they are comfortable with and then try to shoehorn their problem into it. The better approach is to begin with a clearly stated question and then let the question dictate the analysis. Every question needs to be translated into a data problem, and pausing to think through what the results of that analysis would look like and how they might be interpreted prevents you from wasting time on an analysis whose result is uninterpretable.1Leanpub. The Art of Data Science – Section: 3.4 Translating a Question into a Data Problem
A good modeling question is specific enough that you could describe what a useful answer looks like before you run anything. “What affects customer satisfaction?” is vague. “Which of these five service features most strongly predicts whether a customer returns within 90 days?” is a question you can actually model. The specificity forces you to identify your outcome variable, your candidate predictors, and the population you care about. It also helps you figure out how much data you need. For sample size estimation, researchers need to specify the statistical analysis they plan to use, acceptable precision levels, study power, confidence level, and the size of the effect they consider meaningful.2PubMed Central. Sample size determination: A practical guide for health researchers Running a power analysis before collecting data, or at least before committing to a model, saves you from building something that was never going to detect the effect you cared about.
Explore the Data Before Modeling It
Exploratory data analysis is exactly what it sounds like: looking at the data informally before running any formal models. You plot distributions, check for outliers, look at how variables relate to each other, and generally get a feel for what the data contains. This step has two goals: describing what the data actually looks like, and starting to form ideas about which model might be appropriate.3ScienceDirect (Elsevier). Exploratory data analysis
Skipping exploration is tempting, especially when deadlines are tight, but it is one of the fastest ways to end up with a model that looks fine on paper but tells you nothing useful. A quick histogram might reveal that your outcome variable is heavily skewed, which changes your choice of model. A scatter plot might show that the relationship between two variables is curved rather than straight, meaning a simple linear approach will miss the pattern. Exploration also exposes data quality problems early: variables that are mostly missing, values that are clearly wrong, or categories with only a handful of observations.
Clean and Prepare the Data
Real data is messy. Values are missing, formats are inconsistent, and some observations look like they were entered by someone having a very bad day. Data preparation includes fixing or removing errors, standardizing formats, and deciding what to do about missing values. That last item deserves special attention because how you handle missing data can meaningfully change your results.
If you simply delete every row that has a missing value, you may throw away a large portion of your data and introduce bias if the missingness is not random. Imputation, where you fill in missing values based on patterns in the rest of the data, is often a better approach. But the method matters: inappropriate imputation can introduce serious bias in parameter estimates. Research comparing different techniques has found that multiple imputation tends to perform best when less than about 30% of data is missing, producing the highest explained variance and lowest error. Beyond that threshold, even strong imputation methods struggle.4AIP Conference Proceedings. A comparison of model-based imputation methods for handling missing predictor values in a linear regression model: A simulation study If a variable has more than a third of its values missing, you may be better off dropping it entirely or collecting more data.
Other preparation steps depend on your model type. Some models need numeric variables to be scaled to a common range. Categorical variables usually need to be converted into a numeric format the model can digest. And if you plan to split your data into training and testing sets for validation, the split should happen before any transformations that use information from the full dataset, to avoid leaking information from the test set into the training process.
Choose the Right Model Family
The type of model you use depends primarily on what your outcome variable looks like and what question you are asking. If you are trying to predict a continuous number, like revenue or blood pressure, a linear regression family is the natural starting point. If your outcome is binary, like whether a patient survives or a customer churns, logistic regression or a related classification model is appropriate. If you are studying time until an event occurs, survival models come into play. The general form is similar across these families: the response is expressed as a weighted combination of predictors, but the mathematical link between predictors and outcome changes depending on the type of response variable.5ScienceDirect (Elsevier). Regression Modeling Strategies
Beyond the outcome type, the structure of your data influences model choice. If your data includes repeated measurements on the same individuals, such as quarterly health assessments or monthly performance reviews, the observations within each person are not independent. Standard regression treats every row as independent, which can produce misleadingly narrow confidence intervals when the data is actually clustered. Mixed models handle this by accounting for both overall trends and individual-level variation.6Fertility and Sterility. Mixed models for repetitive measurements: the basics Multilevel models serve a similar role, and they are robust against violations of standard assumptions and handle missing data well.7Speech Communication. On multi-level modeling of data from repeated measures designs: a tutorial
Select Which Variables to Include
Deciding which predictors belong in the model is one of the most consequential steps and, honestly, one of the hardest. Including too many variables can make a model overfit the noise in your particular dataset, producing results that look great in training but fall apart on new data. Including too few means you might miss important relationships or leave confounders uncontrolled. Variable selection can improve interpretability and predictive accuracy, but it can also lead to false exclusion of important variables, biased coefficient estimates, and misleading confidence intervals.8PLOS ONE. Evaluating variable selection methods for multivariable regression models: A simulation study protocol
Several automated methods exist. Stepwise selection adds or removes variables one at a time based on statistical criteria. Lasso regression takes a different approach, shrinking less important variable weights toward zero, effectively performing selection and estimation simultaneously. Comparisons between these methods have found that lasso regression tends to be more accurate in selecting the right variables and yields more consistent prediction accuracy than stepwise regression, especially in noisy datasets.9Methodology. Comparison of Lasso and Stepwise Regression in Psychological Data That said, neither approach dominates everywhere. Broader comparisons have found that best subset selection, which considers all possible combinations, performs better when the signal in the data is strong relative to the noise, while lasso performs better when the signal is weak. A variant called the relaxed lasso tends to perform well across both scenarios.10arXiv. Extended Comparisons of Best Subset Selection, Forward Stepwise Selection, and the Lasso
Automated selection should not replace thinking. Subject-matter knowledge about which variables are likely to matter, which are confounders, and which might be mediators or colliders is essential context that no algorithm can supply on its own.
Check Your Assumptions
Every statistical model rests on assumptions, and violating them can make your results unreliable. For linear regression, the main assumptions involve the residuals: the differences between what the model predicts and what actually happened. You want residuals that are roughly normally distributed, have constant spread across the range of predictions, and are independent of each other. Logistic and survival models have their own sets of assumptions.
The standard way to check assumptions is through diagnostic plots and tests, most of which examine residuals in one way or another.11PubMed. Statistical primer: checking model assumptions with regression diagnostics A plot of residuals against predicted values, for example, should show a formless cloud. If it shows a funnel shape, the variance of your errors is not constant and your confidence intervals may be wrong. If it shows a curve, the relationship is not linear and you may need to transform a variable or use a different model.
When assumptions are violated, you have several options: transform a variable, use a different model that relaxes the violated assumption, or use robust methods that remain valid even when assumptions are broken. The worst option is to ignore the violation and proceed as though the assumptions hold, which is unfortunately common.
Evaluate Model Fit
Once a model is fitted, you need to know how well it actually describes the data. Basic fit measures like R-squared, which captures the proportion of variance explained, are a starting point but have limitations. R-squared always increases when you add more predictors, even if those predictors are noise. Adjusted R-squared corrects for this by penalizing model complexity. Information criteria like AIC and BIC take a different approach: they measure how well the model fits relative to how complex it is, with lower values indicating a better balance. BIC penalizes complexity more heavily than AIC, so it tends to favor simpler models.12PPLS Summer Training. Introductory Resources: Statistics and R – Section: Model selection
Which metric to use depends partly on your goal. If you are comparing several candidate models and want to pick the one that best balances accuracy and simplicity, AIC or BIC is typically more informative than raw R-squared. For prediction tasks, metrics like mean squared error on held-out data are usually more relevant than any in-sample statistic. The key principle is that a model that explains training data well is not necessarily a model that will perform well on new data.
Validate on Data the Model Has Not Seen
Validation is about testing whether the patterns your model found are real or just artifacts of the specific dataset you used. The simplest approach is to split your data into a training set and a test set: build the model on the training data and evaluate its performance on the test data. Cross-validation extends this idea by rotating which portion of the data serves as the test set, giving you a more stable estimate of performance.
Cross-validation is widely used, but it comes with subtleties that are easy to miss. Because each data point is used for both training and testing across different folds, the accuracy estimates from each fold are correlated, which means the usual variance estimate is too small. Research has shown that a nested cross-validation approach can produce confidence intervals with better coverage. The same work also found that when constructing confidence intervals from a simple train-test split, you should not refit the model on the combined data afterward, because doing so invalidates those intervals.13PubMed Central. Cross-validation: what does it estimate and how well does it do it?
External validation, where you test the model on an entirely separate dataset collected independently, is the gold standard. It tells you whether the model generalizes beyond the particular time, place, and population it was trained on. Models that perform well internally but have never been externally validated should be treated with caution, especially in high-stakes settings like medicine or criminal justice.
Inference Versus Prediction
One of the biggest sources of confusion in applied modeling is the difference between building a model for inference and building one for prediction. These goals are not the same, and the choices that optimize one often hurt the other.
An inference model is built to understand relationships: does smoking cause lung cancer, does a training program improve test scores, does a new drug lower blood pressure? The focus is on unbiased estimates of the coefficients, proper confidence intervals, and correct interpretation of which variables have real effects. Variable selection in inference models has to be very careful, because aggressive automated selection can bias the remaining coefficients and shrink standard errors to the point where confidence intervals are meaningless.
A prediction model, by contrast, just needs to be right about the outcome. You might not care whether a particular variable’s coefficient is biased, as long as the overall prediction is accurate. This is why techniques like lasso and ridge regression, which deliberately introduce a small amount of bias in exchange for more stable predictions, are popular in prediction contexts but controversial in inference contexts.
The distinction matters practically. If you are using a model to decide whether a variable has a causal effect, you need the inference framework. If you are using a model to forecast next quarter’s sales, the prediction framework is what you want. Conflating them leads to conclusions that sound scientific but are not actually supported by the analysis.
When You Need Causal Claims, Use Causal Tools
Standard regression tells you about associations, not causes. Two variables can be strongly correlated in a model without one causing the other, because a third variable may drive both. If your question is genuinely causal, like “does this treatment improve outcomes,” you need additional structure beyond fitting a regression.
Directed acyclic graphs, or DAGs, are a tool for mapping out the assumed causal relationships among variables before you run any analysis. They help you identify which variables are confounders that need to be controlled for, which are mediators that you might not want to control for, and which are colliders that would actually introduce bias if you adjusted for them.14PubMed Central. Directed acyclic graphs for clinical research: a tutorial This matters because statistical tests alone reveal only the strength of an association between two variables, not the causal relationship, and the researcher must rely on causal reasoning to distinguish confounders from mediators and colliders.15Pediatric Research. Directed acyclic graphs: a tool for causal studies in paediatrics
Drawing a DAG before building your model forces you to make your assumptions explicit. If you assume that variable A affects variable B only through variable C, that assumption is visible on the graph. Other researchers can critique it, and you can test parts of it against data. Skipping this step means your causal assumptions are still there, they are just hidden in the choice of which variables you included and excluded.
Frequentist and Bayesian Approaches
When fitting a model, there are two broad philosophies for estimating its parameters. The frequentist approach, which is what most introductory courses teach, treats the parameters as fixed but unknown quantities and uses the data to estimate them along with confidence intervals. The Bayesian approach treats parameters as uncertain quantities described by probability distributions, combining prior beliefs with the data to produce a posterior distribution that represents updated uncertainty.16Canadian Geotechnical Journal. Statistical characterization of random field parameters using frequentist and Bayesian approaches
For many problems, both approaches give similar answers when you have a reasonable amount of data. The Bayesian approach becomes particularly useful when data is sparse and you have genuine prior information worth incorporating, or when you need to make probability statements directly about the parameters rather than about hypothetical repeated samples. Frequentist methods are computationally simpler and have well-established procedures for most standard models. Choosing between them is less about which is “correct” and more about which aligns with your problem’s needs and your audience’s expectations.
Keeping a Model Alive After Deployment
Building the model is not the end of the process if the model is going to be used in production over time. The world changes, and a model trained on last year’s data may become less accurate as patterns shift. This phenomenon, sometimes called concept drift, is a significant challenge for deployed models. Changes in the underlying system can lead to performance degradation during the model’s lifecycle, and recent work has focused on detecting such changes by monitoring the model’s predictive accuracy over time.17Knowledge-Based Systems. From concept drift to model degradation: An overview on performance-aware drift detectors
Practical monitoring involves tracking key performance metrics on incoming data and flagging when those metrics drop below acceptable thresholds. Some systems retrain automatically on fresh data at regular intervals; others trigger retraining only when degradation is detected. Either way, treating a model as a static artifact that never needs updating is a recipe for decisions based on outdated patterns. Any model that influences real decisions should have a monitoring plan from the start, not as an afterthought.
Fairness and Bias in Statistical Models
A model can be statistically sound and still cause harm if it treats certain groups unfairly. This concern applies to predictive models used in hiring, lending, healthcare, and criminal justice, where biased predictions translate directly into biased outcomes for real people. Many tools now exist to compute fairness metrics, but the field is largely rule-based: if a metric falls within a specified range, the model passes. The problem is that these metrics are statistical estimates based on limited sample data and are therefore subject to sampling variability. If a fairness criterion is barely met or missed, it is often uncertain whether the result is a genuine pass or failure.18AI and Ethics. Bringing practical statistical science to AI and predictive model fairness testing
Beyond the measurement challenge, the choice of which fairness metric to use involves philosophical and legal considerations that statistics alone cannot resolve. Different definitions of fairness can conflict with each other mathematically: a model cannot always satisfy multiple fairness criteria simultaneously, and choosing which criterion matters most requires value judgments about what constitutes fair treatment.19AI and Ethics. The statistical fairness field guide: perspectives from social and formal sciences If your model will affect people’s lives, this is a step that deserves as much rigor as the variable selection or validation steps. Treating fairness as an optional add-on, rather than an integral part of the modeling process, is increasingly untenable both ethically and legally.