πŸ“ˆ

Linear Regression: Statistics Study Notes

October 10, 2026

πŸ“ˆ Comprehensive Guide to Linear Regression

  • Introduction and Core Concepts: Definition, types, and foundational role in statistics and machine learning
  • Model Structure & Categories: Simple vs. multiple linear regression, underlying assumptions, and interpretation of parameters
  • Estimation Methods: Mathematical derivation of least-squares estimation and maximum-likelihood estimation techniques

πŸ’‘ Core Concepts and Overview

In statistics and machine learning, linear regression is a foundational model that estimates the relationship between a scalar response (dependent variable) and one or more explanatory variables (regressor or independent variable) related via a linear combination.

Key Characteristics

  • Supervised Learning Algorithm: Linear regression learns from labeled datasets and maps data points to the most optimized linear functions for predicting new datasets.
  • Probabilistic Focus: It focuses on the conditional probability distribution of the response given the values of predictors, rather than the joint probability distribution of all variables (which falls under multivariate analysis).
  • Historical Significance: It was the first type of regression analysis to be studied rigorously and used extensively because linear models are easier to fit and their statistical properties are simpler to determine than non-linear models.
  • Generalization: A direct generalization of linear regression is found in nonlinear regression.

Primary Practical Use Cases

Linear regression applications generally fall into two broad categories:

  1. Prediction and Forecasting:
    • Fit a predictive model to an observed dataset with the best fit or least error (variance).
    • Use the fitted model to predict response values for new collections of explanatory variables.
  2. Explanation and Quantification:
    • Quantify the strength of the relationship between the response and explanatory variables.
    • Determine whether specific explanatory variables have a linear relationship with the response at all.
    • Identify which subsets of explanatory variables contain redundant information.

πŸ“ Model Categories & Assumptions

Simple vs. Multiple Linear Regression

  • Simple Linear Regression: Involves exactly one explanatory variable (scalar predictor variable xx) and a single scalar response variable yy.
  • Multiple Linear Regression: Involves two or more explanatory variables (denoted with a capital vector Xβƒ—\vec{X}).

Modeling Functions and Error Handling

  • Linear Predictor Functions: Relationships are modeled using linear predictor functions where unknown parameters are estimated from data.
  • Conditional Mean: Most commonly, the conditional mean of the response is assumed to be an affine function of the predictors (less commonly, the conditional median or another quantile is used).
  • Cost Functions and Outliers:
    • Standard models use the Mean Squared Error (MSE) via the least squares approach.
    • Risk with Outliers: MSE assigns higher importance to large errors, meaning datasets with many large outliers can skew the model toward the outliers rather than the true data. Robust cost functions should be used in such cases.

πŸ” Interpretation of Parameters

A fitted linear regression model identifies the relationship between a single predictor variable xjx_j and the response variable yy when all other predictor variables are "held fixed".

Key Interpretive Concepts

  • Unique Effect (Ξ²j\beta_j): The expected change in yy for a one-unit change in xjx_j when other covariates are held fixed (the expected value of the partial derivative of yy with respect to xjx_j).
  • Marginal Effect: Assessed using a correlation coefficient or simple linear regression relating only xjx_j to yy (the total derivative of yy with respect to xjx_j).

Interpretive Caveats and Nuances

  • Non-Marginal Regressors: Care must be taken with regressors that do not allow marginal changes (e.g., dummy variables or the intercept term) or cannot be held fixed simultaneously (like polynomial terms such as ti2t_i^2).
  • Unique vs. Marginal Discrepancies:
    • Near-zero unique effect with large marginal effect: Another covariate captures all information in xjx_j, rendering its unique contribution redundant.
    • Large unique effect with near-zero marginal effect: Other covariates explain a great deal of variation in a complementary way, strengthening the apparent relationship once included.
  • Meaning of "Held Fixed":
    • In a study design (experimental): Literally corresponds to comparisons among units where the experimenter directly sets and holds values constant.
    • In an observational study: Refers to restricting attention to subsets of data that happen to share a common value for the given predictor variable.

βš™οΈ Estimation Procedures

Parameter estimation methods differ in computational simplicity, closed-form solution availability, robustness to heavy-tailed distributions, and theoretical assumptions.

1. Least-Squares Estimation and Related Techniques

Assuming independent variables xiβƒ—=[x1i,x2i,…,xmi]\vec{x_i} = [x_1^i, x_2^i, \ldots, x_m^i] and parameters Ξ²βƒ—=[Ξ²0,Ξ²1,…,Ξ²m]\vec{\beta} = [\beta_0, \beta_1, \ldots, \beta_m], the prediction is: yiβ‰ˆΞ²0+βˆ‘j=1mΞ²jΓ—xjiy_i \approx \beta_0 + \sum_{j=1}^m \beta_j \times x_j^i

By extending xiβƒ—=[1,x1i,x2i,…,xmi]\vec{x_i} = [1, x_1^i, x_2^i, \ldots, x_m^i], the prediction becomes a dot product: yiβ‰ˆβˆ‘j=0mΞ²jΓ—xji=Ξ²βƒ—β‹…xiβƒ—y_i \approx \sum_{j=0}^m \beta_j \times x_j^i = \vec{\beta} \cdot \vec{x_i}

Loss Function & Derivation

The optimum parameter vector minimizes the sum of squared loss: Ξ²^βƒ—=\mboxargminβ⃗ L(D,Ξ²βƒ—)=\mboxargminΞ²βƒ—βˆ‘i=1n(Ξ²βƒ—β‹…xiβƒ—βˆ’yi)2\vec{\hat{\beta}} = {\underset {\vec {\beta }}{\mbox{arg min}}}\,L\left(D,{\vec {\beta }}\right) = {\underset {\vec {\beta }}{\mbox{arg min}}}\sum _{i=1}^{n}\left({\vec {\beta }}\cdot {\vec {x_{i}}}-y_{i}\right)^{2}

Using matrix notation for XX and YY, the loss function expands to: L(D,Ξ²βƒ—)=βˆ₯XΞ²βƒ—βˆ’Yβˆ₯2=(XΞ²βƒ—βˆ’Y)T(XΞ²βƒ—βˆ’Y)=YTYβˆ’YTXΞ²βƒ—βˆ’Ξ²βƒ—TXTY+Ξ²βƒ—TXTXΞ²βƒ—\begin{aligned}L\left(D,{\vec {\beta }}\right)&=\|X{\vec {\beta }}-Y\|^{2}\\&=\left(X{\vec {\beta }}-Y\right)^{\textsf {T}}\left(X{\vec {\beta }}-Y\right)\\&=Y^{\textsf {T}}Y-Y^{\textsf {T}}X{\vec {\beta }}-{\vec {\beta }}^{\textsf {T}}X^{\textsf {T}}Y+{\vec {\beta }}^{\textsf {T}}X^{\textsf {T}}X{\vec {\beta }}\end{aligned}

Setting the gradient of the convex loss function to zero yields: βˆ’2XTY+2XTXΞ²βƒ—=0β‡’XTXΞ²βƒ—=XTYβ‡’Ξ²^βƒ—=(XTX)βˆ’1XTY\begin{aligned}-2X^{\textsf {T}}Y+2X^{\textsf {T}}X{\vec {\beta }}&=0\\\Rightarrow X^{\textsf {T}}X{\vec {\beta }}&=X^{\textsf {T}}Y\\\Rightarrow {\vec {\hat {\beta }}}&=\left(X^{\textsf {T}}X\right)^{-1}X^{\textsf {T}}Y\end{aligned}

Note: To confirm Ξ²^βƒ—\vec{\hat{\beta}} is a minimum, the Hessian matrix must be shown to be positive definite (guaranteed by the Gauss–Markov theorem).

Linear Least Squares Methods Include:

  • Ordinary Least Squares (OLS)
  • Weighted Least Squares (WLS)
  • Generalized Least Squares (GLS)
  • Linear Template Fit

2. Maximum-Likelihood Estimation (MLE)

Maximum likelihood estimation applies when error term distributions belong to a known parametric family fΞΈf_\theta of probability distributions.

  • Normal Distribution Equivalence: When fΞΈf_\theta is a normal distribution with zero mean and variance ΞΈ\theta, the MLE estimate is identical to the OLS estimate.
  • GLS Equivalence: GLS estimates match MLE when errors follow a multivariate normal distribution with a known covariance matrix.

Likelihood Function Formulation

Let data points be (xiβƒ—,yi)({\vec {x_{i}}},y_{i}), parameters be Ξ²βƒ—\vec{\beta}, dataset be DD, and cost function be L(D,Ξ²βƒ—)=βˆ‘i(yiβˆ’Ξ²βƒ—β€‰β‹…β€‰xiβƒ—)2L(D,{\vec {\beta }})=\sum _{i}(y_{i}-{\vec {\beta }}\,\cdot \,{\vec {x_{i}}})^{2}.

Assuming the dependent variable yy follows a Gaussian distribution with fixed standard deviation σ\sigma and a mean that is a linear combination of x⃗\vec{x}:

H(D,Ξ²βƒ—)=∏i=1nPr(yi∣xi⃗  ;Ξ²βƒ—,Οƒ)=∏i=1n12πσexp⁑(βˆ’(yiβˆ’Ξ²βƒ—β€‰β‹…β€‰xiβƒ—)22Οƒ2)\begin{aligned} H(D,{\vec {\beta }})&=\prod _{i=1}^{n}Pr(y_{i}|{\vec {x_{i}}}\,\,;{\vec {\beta }},\sigma )\\ &=\prod _{i=1}^{n}{\frac {1}{{\sqrt {2\pi }}\sigma }}\exp \left(-{\frac {\left(y_{i}-{\vec {\beta }}\,\cdot \,{\vec {x_{i}}}\right)^{2}}{2\sigma ^{2}}}\right) \end{aligned}

Maximizing Log-Likelihood

To simplify optimization, we maximize the strictly increasing logarithmic transformation of the likelihood function (I(D,Ξ²βƒ—)I(D, \vec{\beta})):

I(D,Ξ²βƒ—)=log⁑∏i=1nPr(yi∣xi⃗  ;Ξ²βƒ—,Οƒ)=nlog⁑12Ο€Οƒβˆ’12Οƒ2βˆ‘i=1n(yiβˆ’Ξ²βƒ—β€‰β‹…β€‰xiβƒ—)2\begin{aligned} I(D,{\vec {\beta }})&=\log \prod _{i=1}^{n}Pr(y_{i}|{\vec {x_{i}}}\,\,;{\vec {\beta }},\sigma )\\ &=n\log {\frac {1}{{\sqrt {2\pi }}\sigma }}-{\frac {1}{2\sigma ^{2}}}\sum _{i=1}^{n}\left(y_{i}-{\vec {\beta }}\,\cdot \,{\vec {x_{i}}}\right)^{2} \end{aligned}

The optimal parameter is found via: \mboxargmaxβ⃗ I(D,Ξ²βƒ—)=\mboxargmaxΞ²βƒ—(nlog⁑12Ο€Οƒβˆ’12Οƒ2βˆ‘i=1n(yiβˆ’Ξ²βƒ—β€‰β‹…β€‰xiβƒ—)2){\underset {\vec {\beta }}{\mbox{arg max}}}\,I(D,{\vec {\beta }}) = {\underset {\vec {\beta }}{\mbox{arg max}}}\left(n\log {\frac {1}{{\sqrt {2\pi }}\sigma }}-{\frac {1}{2\sigma ^{2}}}\sum _{i=1}^{n}\left(y_{i}-{\vec {\beta }}\,\cdot \,{\vec {x_{i}}}\right)^{2}\right)


πŸ”¬ Alternative Fitting Approaches

Beyond standard least squares and maximum likelihood, linear models can be fitted using alternative cost functions:

  • Least Absolute Deviations Regression: Minimizes the "lack of fit" using alternative norms.
  • Ridge Regression (L2L_2-norm penalty): Minimizes a penalized version of the least squares cost function to control model complexity.
  • Lasso Regression (L1L_1-norm penalty): Applies an L1L_1-norm penalty, enabling feature selection by driving coefficients to zero.