ML Atlas

04 · Evaluation · 5 min read · Interactive · updated

What is cross-validation and why is it used?

In short

Cross-validation splits the data into k parts and uses each in turn as the validation set. It gives a more stable estimate of performance than a single split.

What it is

Cross-validation (CV) is a method for estimating how well a model will perform on new data: the data is split into k equal parts (folds), and the model is trained k times on k − 1 folds and evaluated on the remaining one. The final result is the average of the k scores, and their spread tells you how much the result depends on the split.

Intuition: a single train/validation split is like grading a student on one pop quiz — it might happen to have easy or hard questions. Cross-validation is five quizzes on different topics: every example lands in the "exam" once, and the average is far less a matter of luck.

The most common choices are k = 5 or k = 10. The extreme case, k equal to the number of examples, is leave-one-out cross-validation.

Mechanism — why it works this way

A score on a single validation set has two sources of uncertainty: the randomness of which examples ended up in validation, and the fact that the validation set is small. Cross-validation tackles both at once. Every example is used for evaluation exactly once, so the estimate rests on all n examples rather than on 20% of them. Averaging k scores reduces the influence of a lucky or unlucky split.

There is also a cost and a subtlety. Each of the k models is trained on (k − 1)/k of the data, i.e. on a slightly smaller set than the final model. That gives CV a slight pessimistic bias — a model trained on all the data will usually be a little better. The larger k, the smaller this bias, but the higher the computational cost and the more similar the training sets become to one another, which can increase the variance of the estimate. Values of 5–10 are an empirical compromise, described among others by Kohavi.

An important caveat: the standard deviation of the k fold scores is not a valid standard error of the mean. The folds share most of their training data, so their scores are correlated, and the simple formula s/√k understates the uncertainty. Treat the spread across folds as a rough signal of stability, not as a confidence interval.

Cross-validation assumes that examples are exchangeable — that every split is equally "fair". When the data has structure (time order, groups of patients, sorted classes), ordinary random folds give inflated or deflated scores, and an appropriate variant must be used. Finally: if CV is used to pick the best of many models, the winner's score is optimistic again — an honest estimate then requires nested cross-validation or a separate test set.

By example

On the Wine dataset (178 wines, 3 cultivars) I compared two procedures for evaluating the same decision tree (DecisionTreeClassifier(random_state=0)). First, 100 single 80/20 splits with random_state from 0 to 99: test accuracy ranged from 75% to 100%, averaging 90.9% with a standard deviation of 4.8 points. Depending on the draw, the same model could be described as "mediocre" or "flawless".

Next, 100 repetitions of stratified 5-fold cross-validation (StratifiedKFold(5, shuffle=True) with different seeds): the five-fold averages now ranged only from 86.5% to 94.4%, with a standard deviation of 1.6 points — three times less. The mean estimate stayed the same (about 90.9%); only most of the randomness disappeared. For a single run with random_state=0 the fold scores were 91.7%, 83.3%, 97.2%, 97.1% and 94.3% (mean 92.7%) — showing how much the folds themselves differ.

In practice

  • The simplest call: cross_val_score(model, X, y, cv=5); for classification scikit-learn uses StratifiedKFold by default.
  • cross_validate(..., scoring=['accuracy', 'f1_macro'], return_train_score=True) returns several metrics at once, plus the training score for diagnosing overfitting.
  • Always pass the whole Pipeline (scaling, imputation, model), not pre-processed X — otherwise statistics from the validation folds leak into training.
  • When the data might be sorted, set shuffle=True and random_state; for groups use GroupKFold, for time series TimeSeriesSplit.
  • With very small datasets, consider RepeatedStratifiedKFold to average out the randomness of the fold assignment as well.
  • Typical mistake: selecting features on the full data and only then running CV — the score can be inflated by tens of points.

Frequently asked questions

Which k should I choose?
5 or 10 by default. With large datasets 3–5 is enough, since each fold is large anyway and training cost grows linearly with k. With very small datasets, repeated 5- or 10-fold CV is better than leave-one-out, which can be unstable.
Which of the k models should I pick at the end?
None of them. Cross-validation evaluates the training procedure, not a specific instance of the model. Once the settings are chosen, a single model is trained on all the available training data.
Does cross-validation replace the test set?
For estimating the quality of a single model fixed in advance — yes. When you use CV for tuning or for choosing among many models, the best one's score is optimistic, and you need a separate test set or nested CV.

Sources

  • James G., Witten D., Hastie T., Tibshirani R. "An Introduction to Statistical Learning", 2nd ed., Springer 2021, ch. 5.1 (Cross-Validation).
  • Hastie T., Tibshirani R., Friedman J. "The Elements of Statistical Learning", 2nd ed., Springer 2009, ch. 7.10 (Cross-Validation).
  • Stone M. (1974). Cross-Validatory Choice and Assessment of Statistical Predictions. Journal of the Royal Statistical Society, Series B, 36(2).
  • Kohavi R. (1995). A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection. Proceedings of IJCAI 1995.
  • scikit-learn documentation: Cross-validation: evaluating estimator performance — https://scikit-learn.org/stable/modules/cross_validation.html

See also