Central Limit Theorem: Statistics Study Notes
October 10, 2026
📚 The Central Limit Theorem (CLT): A Comprehensive Guide
- Core Overview and Fundamental Concepts
- Historical Context and Development
- Independent Sequences and Classical Variants
- Classical CLT
- Lyapunov CLT
- Lindeberg (-Feller) CLT
- Specialized CLT Extensions
- CLT for Random Number of Variables
- Multidimensional (Multivariate) CLT
- Generalized Central Limit Theorem (GCLT)
- Dependent Processes and Weak Dependence
💡 Fundamental Concepts of the Central Limit Theorem
In probability theory, the central limit theorem (CLT) states that, under appropriate conditions, the distribution of a normalized version of the sample mean converges to a standard normal distribution.
The theorem is a key concept in probability theory because it implies that probabilistic and statistical methods that work for normal distributions can be applicable to many problems involving other types of distributions.
🔑 Key Characteristics
- Distribution Independence: This holds even if the original variables themselves are not normally distributed.
- Large Sample Approximation: When a large sample of independent observations is obtained and averaged repeatedly, the probability distribution of these averages will closely approximate a normal distribution if the sample size is large enough.
- Underlying Variations: There are several versions of the CLT, each applying in the context of different conditions (e.g., independent and identically distributed vs. varying conditions).
📐 Statistical Formulation
Let denote a statistical sample of size from a population with:
- Expected value (average):
- Finite positive variance:
- Sample mean: (itself a random variable)
The limit as of the distribution of is a normal distribution with:
- Mean:
- Variance:
⏳ Historical Development
The central limit theorem has seen many changes during the formal development of probability theory:
- 1811: Previous versions of the theorem date back to this era.
- De Moivre–Laplace Theorem: The earliest version, stating that the normal distribution may be used as an approximation to the binomial distribution.
- 1920s: The theorem was stated in its precise modern form.
🔬 Independent Sequences & Classical Variants
1. Classical CLT
Let be a sequence of i.i.d. (independent and identically distributed) random variables having:
- Expected value:
- Finite variance:
The sample average is defined as:
📊 Convergence Mechanics
- Law of Large Numbers: The sample average converges almost surely (and in probability) to the expected value as .
- Stochastic Fluctuations: The classical CLT describes the size and distributional form of the stochastic fluctuations around during this convergence.
- Normalized Mean: The scaled difference approaches the normal distribution with mean and variance .
- Distribution Proximity: For large enough , the distribution of gets arbitrarily close to the normal distribution with mean and variance .
🧮 Formal Mathematical Statement
For , convergence in distribution means cumulative distribution functions converge pointwise to the cdf of the distribution. For every real number :
Where:
- is the standard normal cdf evaluated at .
- The convergence is uniform in :
- denotes the supremum (least upper bound) of the set.
2. Lyapunov CLT
- Independence Requirement: The random variables must be independent, but not necessarily identically distributed.
- Moments Condition: Random variables must have moments of some order .
- Growth Rate: The rate of growth of these moments is limited by the Lyapunov condition.
- Practical Tip: In practice, it is usually easiest to check Lyapunov's condition for .
- Relationship: If a sequence of random variables satisfies Lyapunov's condition, it also satisfies Lindeberg's condition (though the converse does not hold).
3. Lindeberg (-Feller) CLT
Operating in the same setting as the classical model, Lindeberg (1920) replaced the Lyapunov condition with a weaker alternative.
📋 Lindeberg Condition
Suppose that for every : (where is the indicator function).
🎯 Result
The distribution of the standardized sums: converges towards the standard normal distribution .
🌐 Specialized CLT Extensions
1. CLT for a Random Number of Variables
Rather than summing a fixed integer number of random variables and taking , the sum can consist of a random number of random variables, subject to specific conditions on .
- For example, Robbins' Theorem (1948, Corollary 4) assumes that is asymptotically normal (alongside other derived conditions leading to the same result).
2. Multidimensional (Multivariate) CLT
Proofs using characteristic functions extend to cases where each individual is a random vector in , featuring:
- Mean vector:
- Covariance matrix:
- Properties: Independent and identically distributed random vectors.
🧮 Vector Definitions
For :
📈 Multivariate Statement
The multivariate central limit theorem states that:
Where the covariance matrix equals:
- Proof Methodology: Proved using the Cramér–Wold theorem.
- Convergence Rate: Given by a Berry–Esseen type result (though whether the factor is necessary remains unknown).
🔄 The Generalized Central Limit Theorem (GCLT)
The generalized central limit theorem was developed through the collaborative efforts of multiple mathematicians—including Sergei Bernstein, Jarl Waldemar Lindeberg, Paul Lévy, William Feller, and Andrey Kolmogorov—between 1920 and 1937.
- First Complete Proof: Published in 1937 by Paul Lévy in French.
- English Availability: Found in the translated 1954 book by Boris Vladimirovich Gnedenko and Kolmogorov.
- Core Statement: If sums of independent, identically distributed random variables converge in distribution to some , then must be a stable distribution.
🔗 Dependent Processes
CLT Under Weak Dependence
A useful generalization of independent, identically distributed variables is a mixing random process in discrete time.
- Definition: "Mixing" means, roughly, that random variables temporally far apart from one another are nearly independent.
- Applications: Utilized heavily in ergodic theory and probability theory.
- Key Metric: Strong mixing (or -mixing), defined by , where represents the strong mixing coefficient.
📝 Simplified Formulation
Under strong mixing conditions, the variance takes the form: