ML Atlas

11 · Laws · 5 min read · Interactive · updated

What is data leakage in machine learning and how do you avoid it?

In short

Data leakage is when a model, in training or validation, gets information that would not be available at real prediction time. It inflates the results.

What it is

Any information that would not be available at the moment of a real prediction, yet found its way into a model's training or evaluation, inflates its measured performance. The concept was systematised by Shachar Kaufman, Saharon Rosset, Claudia Perlich and Ori Stitelman in 2012; in 2023 Sayash Kapoor and Arvind Narayanan showed that leakage is responsible for errors in hundreds of scientific papers that use ML.

Leakage takes two main forms. Target leakage: a feature contains a trace of the label, e.g. "hospital discharge date" in a model predicting hospitalisation, or "refund amount" in a model predicting complaints. Split leakage: information from the validation or test set seeps into training — through scaling, feature selection, imputation or hyperparameter tuning done on the full data, through duplicates, or through the same entity appearing on both sides of the split.

The symptom is always the same: the validation score is clearly better than the score after deployment. Leakage is the most common reason for "miraculous" results that do not hold up in practice.

Mechanism — why it works this way

Validation is meant to simulate the future: the model is evaluated on data it has not seen. Every processing step fitted to the data — the mean used for imputation, the scaling parameters, the choice of best features — is part of the model. If that step saw the validation data, the model has indirectly learned from it.

Feature selection is the most dangerous. Among thousands of pure-noise features, some will correlate with the label by chance in a given dataset. If we select them on the full data and then run cross-validation, those chance correlations are also present in the validation folds — the model "discovers" a relationship that we ourselves put there. Hastie, Tibshirani and Friedman describe this as "the wrong and the right way to do cross-validation"; in 2002 Ambroise and McLachlan showed that this mistake had inflated the results of papers on classification from gene microarray data.

The strength of leakage depends on the ratio of features to examples and on how much information a given step carries over. Scaling carries little (two numbers per feature), so with large datasets its effect can be negligible. Selection from thousands of features carries a great deal. Target leakage can produce a near-perfect score regardless of everything else.

Time series add leakage from the future: a random split mixes past and future, and features computed over windows (moving averages) may contain values from after the moment of prediction.

By example

Breast Cancer Wisconsin (569 cases, 30 features), 5-fold cross-validation. Standardisation before the split, then kNN: 0.965. Standardisation inside a pipeline: also 0.965. Selecting the 5 best features from 30 real and 2,000 noise features before the split: 0.947; inside a pipeline: 0.947. On this dataset, with its strong signal, leakage changed nothing — and that is the trap, because it teaches you that "there is no problem".

Now the same labels, but only pure-noise features. Choosing the 20 "best" of 10,000 noise columns before validation gives an accuracy of 0.71 — above the share of the majority class (0.63), so it looks like signal. The same selection inside a pipeline gives 0.50: there is nothing there. On a random subsample of 60 patients the difference is dramatic: 0.98 with leakage versus 0.48 without. A model built on pure noise looked almost perfect.

In practice

  • Put every step that is fitted to data into a Pipeline / make_pipeline together with the model, and only then pass the whole thing to cross_val_score or GridSearchCV.
  • When one entity (patient, customer, session) has many rows, split the data by group: GroupKFold, GroupShuffleSplit.
  • For time series use TimeSeriesSplit and compute features only from the past relative to the moment of prediction.
  • For every feature ask: would I know this value at the moment the decision is made? Check suspiciously important features first.
  • Remove duplicates and near-duplicates before splitting; check that the test set contains no rows from training.

Frequently asked questions

Is scaling before the split really leakage?
Formally yes, because the scaling parameters contain information from the validation data. In practice, with large datasets the effect can be negligible. Even so, it is worth always using a pipeline — the same habit protects you from the dangerous cases, such as feature selection or target encoding.
How do you recognise leakage when the result is simply good?
Warning signs: a score much better than in the literature or than competitors', a single feature dominating the importance ranking, a big drop between validation and new data. In that case, review the origin and creation time of every important feature.
How does leakage differ from shortcut learning?
Leakage is a mistake in data preparation or in the evaluation procedure. Shortcut learning is the behaviour of a model that exploits an incidental feature present in the data. Leakage often creates shortcuts, but a shortcut can exist without leakage.

Sources

  • Kaufman S., Rosset S., Perlich C., Stitelman O. (2012). Leakage in Data Mining: Formulation, Detection, and Avoidance. ACM Transactions on Knowledge Discovery from Data, 6(4), 15.
  • Kapoor S., Narayanan A. (2023). Leakage and the Reproducibility Crisis in Machine-Learning-Based Science. Patterns, 4(9), 100804.
  • Ambroise C., McLachlan G. J. (2002). Selection Bias in Gene Extraction on the Basis of Microarray Gene-Expression Data. PNAS, 99(10), 6562–6566.
  • Hastie T., Tibshirani R., Friedman J. (2009). The Elements of Statistical Learning, 2nd ed. Springer, ch. 7.10.2.
  • scikit-learn: Common pitfalls and recommended practices, https://scikit-learn.org/stable/common_pitfalls.html

See also