📊

Central Limit Theorem: Statistics Study Notes

October 10, 2026

📚 The Central Limit Theorem (CLT): A Comprehensive Guide

  • Core Overview and Fundamental Concepts
  • Historical Context and Development
  • Independent Sequences and Classical Variants
    • Classical CLT
    • Lyapunov CLT
    • Lindeberg (-Feller) CLT
  • Specialized CLT Extensions
    • CLT for Random Number of Variables
    • Multidimensional (Multivariate) CLT
  • Generalized Central Limit Theorem (GCLT)
  • Dependent Processes and Weak Dependence

💡 Fundamental Concepts of the Central Limit Theorem

In probability theory, the central limit theorem (CLT) states that, under appropriate conditions, the distribution of a normalized version of the sample mean converges to a standard normal distribution.

The theorem is a key concept in probability theory because it implies that probabilistic and statistical methods that work for normal distributions can be applicable to many problems involving other types of distributions.

🔑 Key Characteristics

  • Distribution Independence: This holds even if the original variables themselves are not normally distributed.
  • Large Sample Approximation: When a large sample of independent observations is obtained and averaged repeatedly, the probability distribution of these averages will closely approximate a normal distribution if the sample size is large enough.
  • Underlying Variations: There are several versions of the CLT, each applying in the context of different conditions (e.g., independent and identically distributed vs. varying conditions).

📐 Statistical Formulation

Let X1,X2,…,XnX_1, X_2, \dots, X_n denote a statistical sample of size nn from a population with:

  • Expected value (average): μ\mu
  • Finite positive variance: σ2\sigma^2
  • Sample mean: Xˉn\bar{X}_n (itself a random variable)

The limit as n→∞n \to \infty of the distribution of (Xˉn−μ)n({\bar {X}}_{n}-\mu ){\sqrt {n}} is a normal distribution with:

  • Mean: 00
  • Variance: σ2\sigma^2

⏳ Historical Development

The central limit theorem has seen many changes during the formal development of probability theory:

  • 1811: Previous versions of the theorem date back to this era.
  • De Moivre–Laplace Theorem: The earliest version, stating that the normal distribution may be used as an approximation to the binomial distribution.
  • 1920s: The theorem was stated in its precise modern form.

🔬 Independent Sequences & Classical Variants

1. Classical CLT

Let (Xn)n≥1(X_n)_{n\geq 1} be a sequence of i.i.d. (independent and identically distributed) random variables having:

  • Expected value: μ\mu
  • Finite variance: σ2\sigma^2

The sample average is defined as: Xˉn≡X1+⋯+Xnn{\bar {X}}_{n}\equiv {\frac {X_{1}+\cdots +X_{n}}{n}}

📊 Convergence Mechanics

  • Law of Large Numbers: The sample average converges almost surely (and in probability) to the expected value μ\mu as n→∞n \to \infty.
  • Stochastic Fluctuations: The classical CLT describes the size and distributional form of the stochastic fluctuations around μ\mu during this convergence.
  • Normalized Mean: The scaled difference n(Xˉn−μ){\sqrt {n}}({\bar {X}}_{n}-\mu ) approaches the normal distribution with mean 00 and variance σ2\sigma^2.
  • Distribution Proximity: For large enough nn, the distribution of Xˉn{\bar {X}}_{n} gets arbitrarily close to the normal distribution with mean μ\mu and variance σ2/n\sigma^2/n.

🧮 Formal Mathematical Statement

For σ>0\sigma > 0, convergence in distribution means cumulative distribution functions converge pointwise to the cdf of the N(0,σ2){\mathcal {N}}(0,\sigma^2) distribution. For every real number zz:

lim⁡n→∞P[n(Xˉn−μ)≤z]=lim⁡n→∞P[n(Xˉn−μ)σ≤zσ]=Φ(zσ)\lim _{n\to \infty }\mathbb {P} \left[{\sqrt {n}}({\bar {X}}_{n}-\mu )\leq z\right]=\lim _{n\to \infty }\mathbb {P} \left[{\frac {{\sqrt {n}}({\bar {X}}_{n}-\mu )}{\sigma }}\leq {\frac {z}{\sigma }}\right]=\Phi \left({\frac {z}{\sigma }}\right)

Where:

  • Φ(z)\Phi(z) is the standard normal cdf evaluated at zz.
  • The convergence is uniform in zz: lim⁡n→∞  sup⁡z∈R  ∣P[n(Xˉn−μ)≤z]−Φ(zσ)∣=0 \lim _{n\to \infty }\;\sup _{z\in \mathbb {R} }\;\left|\mathbb {P} \left[{\sqrt {n}}({\bar {X}}_{n}-\mu )\leq z\right]-\Phi \left({\frac {z}{\sigma }}\right)\right|=0~
  • sup⁡\sup denotes the supremum (least upper bound) of the set.

2. Lyapunov CLT

  • Independence Requirement: The random variables XiX_i must be independent, but not necessarily identically distributed.
  • Moments Condition: Random variables ∣Xi∣|X_i| must have moments of some order (2+δ)(2+\delta).
  • Growth Rate: The rate of growth of these moments is limited by the Lyapunov condition.
  • Practical Tip: In practice, it is usually easiest to check Lyapunov's condition for δ=1\delta = 1.
  • Relationship: If a sequence of random variables satisfies Lyapunov's condition, it also satisfies Lindeberg's condition (though the converse does not hold).

3. Lindeberg (-Feller) CLT

Operating in the same setting as the classical model, Lindeberg (1920) replaced the Lyapunov condition with a weaker alternative.

📋 Lindeberg Condition

Suppose that for every ε>0\varepsilon > 0: lim⁡n→∞1sn2∑i=1nE⁡[(Xi−μi)2⋅1{∣Xi−μi∣>εsn}]=0\lim _{n\to \infty }{\frac {1}{s_{n}^{2}}}\sum _{i=1}^{n}\operatorname {E} \left[(X_{i}-\mu _{i})^{2}\cdot \mathbf {1} _{\left\{\left|X_{i}-\mu _{i}\right|>\varepsilon s_{n}\right\}}\right]=0 (where 1{…}\mathbf{1}_{\{\ldots\}} is the indicator function).

🎯 Result

The distribution of the standardized sums: 1sn∑i=1n(Xi−μi){\frac {1}{s_{n}}}\sum _{i=1}^{n}\left(X_{i}-\mu _{i}\right) converges towards the standard normal distribution N(0,1){\mathcal {N}}(0,1).


🌐 Specialized CLT Extensions

1. CLT for a Random Number of Variables

Rather than summing a fixed integer number nn of random variables and taking n→∞n \to \infty, the sum can consist of a random number NN of random variables, subject to specific conditions on NN.

  • For example, Robbins' Theorem (1948, Corollary 4) assumes that NN is asymptotically normal (alongside other derived conditions leading to the same result).

2. Multidimensional (Multivariate) CLT

Proofs using characteristic functions extend to cases where each individual Xi\mathbf{X}_i is a random vector in Rk\mathbb{R}^k, featuring:

  • Mean vector: μ=E⁡[Xi]{\boldsymbol {\mu }}=\operatorname {E} [\mathbf {X} _{i}]
  • Covariance matrix: Σ\mathbf {\Sigma }
  • Properties: Independent and identically distributed random vectors.

🧮 Vector Definitions

For i=1,2,3,…i = 1, 2, 3, \dots: Xi=[Xi(1)⋮Xi(k)],∑i=1nXi=[∑i=1nXi(1)⋮∑i=1nXi(k)],Xˉn=1n∑i=1nXi\mathbf{X}_{i}={\begin{bmatrix}X_{i}^{(1)}\\\vdots \\X_{i}^{(k)}\end{bmatrix}}, \quad \sum _{i=1}^{n}\mathbf {X} _{i}={\begin{bmatrix}\sum _{i=1}^{n}X_{i}^{(1)}\\\vdots \\\sum _{i=1}^{n}X_{i}^{(k)}\end{bmatrix}}, \quad \mathbf {{\bar {X}}_{n}} ={\frac {1}{n}}\sum _{i=1}^{n}\mathbf {X} _{i}

📈 Multivariate Statement

The multivariate central limit theorem states that: n(X‾n−μ)⟶dNk(0,Σ){\sqrt {n}}\left({\overline {\mathbf {X} }}_{n}-{\boldsymbol {\mu }}\right)\mathrel {\overset {d}{\longrightarrow }} {\mathcal {N}}_{k}(0,{\boldsymbol {\Sigma }})

Where the covariance matrix Σ\boldsymbol{\Sigma} equals: Σ=[Var⁡(X1(1))Cov⁡(X1(1),X1(2))⋯Cov⁡(X1(1),X1(k))Cov⁡(X1(2),X1(1))Var⁡(X1(2))⋯Cov⁡(X1(2),X1(k))⋮⋮⋱⋮Cov⁡(X1(k),X1(1))Cov⁡(X1(k),X1(2))⋯Var⁡(X1(k))]\boldsymbol {\Sigma }={\begin{bmatrix}{\operatorname {Var} \left(X_{1}^{(1)}\right)}&\operatorname {Cov} \left(X_{1}^{(1)},X_{1}^{(2)}\right)&\cdots &\operatorname {Cov} \left(X_{1}^{(1)},X_{1}^{(k)}\right)\\\operatorname {Cov} \left(X_{1}^{(2)},X_{1}^{(1)}\right)&\operatorname {Var} \left(X_{1}^{(2)}\right)&\cdots &\operatorname {Cov} \left(X_{1}^{(2)},X_{1}^{(k)}\right)\\\vdots &\vdots &\ddots &\vdots \\\operatorname {Cov} \left(X_{1}^{(k)},X_{1}^{(1)}\right)&\operatorname {Cov} \left(X_{1}^{(k)},X_{1}^{(2)}\right)&\cdots &\operatorname {Var} \left(X_{1}^{(k)}\right)\end{bmatrix}}

  • Proof Methodology: Proved using the Cramér–Wold theorem.
  • Convergence Rate: Given by a Berry–Esseen type result (though whether the factor d1/4d^{1/4} is necessary remains unknown).

🔄 The Generalized Central Limit Theorem (GCLT)

The generalized central limit theorem was developed through the collaborative efforts of multiple mathematicians—including Sergei Bernstein, Jarl Waldemar Lindeberg, Paul Lévy, William Feller, and Andrey Kolmogorov—between 1920 and 1937.

  • First Complete Proof: Published in 1937 by Paul Lévy in French.
  • English Availability: Found in the translated 1954 book by Boris Vladimirovich Gnedenko and Kolmogorov.
  • Core Statement: If sums of independent, identically distributed random variables converge in distribution to some ZZ, then ZZ must be a stable distribution.

🔗 Dependent Processes

CLT Under Weak Dependence

A useful generalization of independent, identically distributed variables is a mixing random process in discrete time.

  • Definition: "Mixing" means, roughly, that random variables temporally far apart from one another are nearly independent.
  • Applications: Utilized heavily in ergodic theory and probability theory.
  • Key Metric: Strong mixing (or α\alpha-mixing), defined by α(n)→0\alpha(n) \to 0, where α(n)\alpha(n) represents the strong mixing coefficient.

📝 Simplified Formulation

Under strong mixing conditions, the variance takes the form: σ2=E⁡(X12)+2…\sigma^{2}=\operatorname{E}(X_{1}^{2})+2\dots