Linear Regression: Statistics Study Notes
October 10, 2026
π Comprehensive Guide to Linear Regression
- Introduction and Core Concepts: Definition, types, and foundational role in statistics and machine learning
- Model Structure & Categories: Simple vs. multiple linear regression, underlying assumptions, and interpretation of parameters
- Estimation Methods: Mathematical derivation of least-squares estimation and maximum-likelihood estimation techniques
π‘ Core Concepts and Overview
In statistics and machine learning, linear regression is a foundational model that estimates the relationship between a scalar response (dependent variable) and one or more explanatory variables (regressor or independent variable) related via a linear combination.
Key Characteristics
- Supervised Learning Algorithm: Linear regression learns from labeled datasets and maps data points to the most optimized linear functions for predicting new datasets.
- Probabilistic Focus: It focuses on the conditional probability distribution of the response given the values of predictors, rather than the joint probability distribution of all variables (which falls under multivariate analysis).
- Historical Significance: It was the first type of regression analysis to be studied rigorously and used extensively because linear models are easier to fit and their statistical properties are simpler to determine than non-linear models.
- Generalization: A direct generalization of linear regression is found in nonlinear regression.
Primary Practical Use Cases
Linear regression applications generally fall into two broad categories:
- Prediction and Forecasting:
- Fit a predictive model to an observed dataset with the best fit or least error (variance).
- Use the fitted model to predict response values for new collections of explanatory variables.
- Explanation and Quantification:
- Quantify the strength of the relationship between the response and explanatory variables.
- Determine whether specific explanatory variables have a linear relationship with the response at all.
- Identify which subsets of explanatory variables contain redundant information.
π Model Categories & Assumptions
Simple vs. Multiple Linear Regression
- Simple Linear Regression: Involves exactly one explanatory variable (scalar predictor variable ) and a single scalar response variable .
- Multiple Linear Regression: Involves two or more explanatory variables (denoted with a capital vector ).
Modeling Functions and Error Handling
- Linear Predictor Functions: Relationships are modeled using linear predictor functions where unknown parameters are estimated from data.
- Conditional Mean: Most commonly, the conditional mean of the response is assumed to be an affine function of the predictors (less commonly, the conditional median or another quantile is used).
- Cost Functions and Outliers:
- Standard models use the Mean Squared Error (MSE) via the least squares approach.
- Risk with Outliers: MSE assigns higher importance to large errors, meaning datasets with many large outliers can skew the model toward the outliers rather than the true data. Robust cost functions should be used in such cases.
π Interpretation of Parameters
A fitted linear regression model identifies the relationship between a single predictor variable and the response variable when all other predictor variables are "held fixed".
Key Interpretive Concepts
- Unique Effect (): The expected change in for a one-unit change in when other covariates are held fixed (the expected value of the partial derivative of with respect to ).
- Marginal Effect: Assessed using a correlation coefficient or simple linear regression relating only to (the total derivative of with respect to ).
Interpretive Caveats and Nuances
- Non-Marginal Regressors: Care must be taken with regressors that do not allow marginal changes (e.g., dummy variables or the intercept term) or cannot be held fixed simultaneously (like polynomial terms such as ).
- Unique vs. Marginal Discrepancies:
- Near-zero unique effect with large marginal effect: Another covariate captures all information in , rendering its unique contribution redundant.
- Large unique effect with near-zero marginal effect: Other covariates explain a great deal of variation in a complementary way, strengthening the apparent relationship once included.
- Meaning of "Held Fixed":
- In a study design (experimental): Literally corresponds to comparisons among units where the experimenter directly sets and holds values constant.
- In an observational study: Refers to restricting attention to subsets of data that happen to share a common value for the given predictor variable.
βοΈ Estimation Procedures
Parameter estimation methods differ in computational simplicity, closed-form solution availability, robustness to heavy-tailed distributions, and theoretical assumptions.
1. Least-Squares Estimation and Related Techniques
Assuming independent variables and parameters , the prediction is:
By extending , the prediction becomes a dot product:
Loss Function & Derivation
The optimum parameter vector minimizes the sum of squared loss:
Using matrix notation for and , the loss function expands to:
Setting the gradient of the convex loss function to zero yields:
Note: To confirm is a minimum, the Hessian matrix must be shown to be positive definite (guaranteed by the GaussβMarkov theorem).
Linear Least Squares Methods Include:
- Ordinary Least Squares (OLS)
- Weighted Least Squares (WLS)
- Generalized Least Squares (GLS)
- Linear Template Fit
2. Maximum-Likelihood Estimation (MLE)
Maximum likelihood estimation applies when error term distributions belong to a known parametric family of probability distributions.
- Normal Distribution Equivalence: When is a normal distribution with zero mean and variance , the MLE estimate is identical to the OLS estimate.
- GLS Equivalence: GLS estimates match MLE when errors follow a multivariate normal distribution with a known covariance matrix.
Likelihood Function Formulation
Let data points be , parameters be , dataset be , and cost function be .
Assuming the dependent variable follows a Gaussian distribution with fixed standard deviation and a mean that is a linear combination of :
Maximizing Log-Likelihood
To simplify optimization, we maximize the strictly increasing logarithmic transformation of the likelihood function ():
The optimal parameter is found via:
π¬ Alternative Fitting Approaches
Beyond standard least squares and maximum likelihood, linear models can be fitted using alternative cost functions:
- Least Absolute Deviations Regression: Minimizes the "lack of fit" using alternative norms.
- Ridge Regression (-norm penalty): Minimizes a penalized version of the least squares cost function to control model complexity.
- Lasso Regression (-norm penalty): Applies an -norm penalty, enabling feature selection by driving coefficients to zero.