ML Atlas

11 · Laws · 4 min read · Interactive · updated

What does the central limit theorem say and why is the normal distribution everywhere?

In short

The sum or mean of many independent values is approximately normally distributed, even when the individual values are extremely skewed. Here is why.

What it is

The mean (or sum) of many independent random variables with finite variance is approximately normally distributed with standard deviation σ/√n — regardless of what the distribution of a single variable looks like. The case of a coin was described by Abraham de Moivre in The Doctrine of Chances (2nd ed., 1738), and the general version was given by Pierre-Simon Laplace in 1810.

This explains why the bell curve turns up so often: height, measurement error or a test score are the result of summing many small, independent influences. It also explains why confidence intervals and statistical tests work on data that is not itself normal.

The theorem does not say that the data is normal. It says that the sample mean is normal when the sample is large enough.

Mechanism — why it works this way

Intuition: when you add up many independent components, their shapes "blur together". The skew of one component is partly offset by the others, because extreme values rarely all occur at once. Every addition of random variables is a convolution of distributions, and repeated convolution smooths any shape towards a bell.

Formally: the normal distribution is the only distribution (with finite variance) that is stable under addition — the sum of two normal variables is again normal, with the same shape. It is the fixed point towards which averaging "pulls". In the proof via characteristic functions, all higher moments (skewness, kurtosis) vanish in the limit, leaving only the mean and the variance.

Numbers worth knowing: the skewness of a mean of n observations shrinks like 1/√n. A distribution with a skewness of 2 (the exponential) has a skewness of about 0.37 after averaging 30 values — already nearly symmetric. A distribution with a skewness of 5 needs much larger samples.

Caveats. Finite variance is required: for the Cauchy distribution the mean never becomes normal. Independence (or weak dependence) is required. The approximation is worst in the tails — probabilities of extreme events estimated from the normal distribution can be misleading even for large n. The popular rule "n ≥ 30 is enough" is not a law, just a heuristic for moderately skewed data.

By example

Titanic ticket fares (891 passengers) are extremely skewed: mean £32.2, median £14.5, skewness 4.78. We draw 50 tickets 10,000 times (with replacement, seed 7) and compute the mean of each sample. The distribution of the means has a standard deviation of 7.00 — the theoretical σ/√n gives 7.02 — and the skewness drops from 4.78 to 0.67. Within ±1.96 standard errors of the true mean lie 95.5% of the sample means, almost exactly the predicted 95%.

The same experiment on an exponential distribution (mean 1, skewness 2, seed 7): the skewness of the means is 2.0 for n = 1, 1.45 for n = 2, 0.91 for n = 5, 0.36 for n = 30 and 0.22 for n = 100. The standard deviations of the means (0.70, 0.45, 0.18, 0.099) agree with 1/√n to within hundredths.

In practice

  • Standard error of the mean: x.std(ddof=1) / np.sqrt(len(x)); a 95% confidence interval is the mean ± 1.96 × standard error (in code, scipy.stats.norm.ppf(0.975)).
  • For small samples use the t distribution: scipy.stats.t.interval(0.95, df=n-1, loc=m, scale=se).
  • For very skewed data or unusual statistics (median, AUC, F1), use the bootstrap instead of a formula — scipy.stats.bootstrap.
  • Assess the accuracy difference between two models on the same test set using per-example differences, not two separate intervals.
  • Check the shape of the means by simulation (np.random.default_rng) rather than trusting the "n = 30" rule.
  • Typical mistake: concluding "the data is normal because of the CLT". The CLT is about means, not individual observations; model residuals and outliers have to be inspected separately.

Frequently asked questions

Does the CLT mean my data is normally distributed?
No. The data can be arbitrarily skewed. Only the distribution of the mean (or sum) computed from many independent observations becomes normal.
From what n does the central limit theorem kick in?
It depends on the skewness and tails of the distribution. For symmetric data a few observations suffice, for exponential data about 30, and for prices or incomes several hundred may be needed. The safest check is simulation or the bootstrap.
How does the CLT relate to machine learning?
It justifies confidence intervals for metrics on a test set, explains why mini-batch gradient noise is approximately normal, and underlies many tests for comparing models.

Sources

  • Abraham de Moivre, "The Doctrine of Chances", 2nd ed., London 1738.
  • Pierre-Simon Laplace, memoir on approximations of functions of very large numbers, Mémoires de l'Académie des sciences, Paris 1810.
  • Charles M. Grinstead, J. Laurie Snell, "Introduction to Probability", American Mathematical Society, 1997, ch. 9 (Central Limit Theorem).
  • Larry Wasserman, "All of Statistics", Springer, 2004, ch. 5 (Convergence of Random Variables).

See also