Most bugs announce themselves — a crash, a stack trace, a wrong answer you can see. The most dangerous bug in machine learning does the opposite: it makes everything look better than it is. Your model posts a spectacular score on the test set, everyone is delighted, it ships — and in production it performs like a coin flip. The culprit is almost always data leakage: information that would not be available at prediction time sneaking into training, so the evaluation measures something impossible. It is worth understanding deeply, because it is common, it is silent, and it has produced a documented crisis of over-optimistic results across ML-based research.
Why leakage is so dangerous
The whole point of a test set is to simulate the unseen future: data the model has never met, standing in for the real world it will face. Leakage breaks that simulation by letting the model peek at information it will not actually have when it makes real predictions. The score you get is then an answer to a question you did not ask — "how well does the model do when it can cheat?" — and the gap between that and reality only reveals itself after deployment, when the stakes are highest.
Data leakage
Any situation where the training process has access to information that would not be available at the moment of a real prediction. The model learns to rely on it, the evaluation rewards that reliance, and production — where the information is absent — exposes the illusion.
Form 1: preprocessing before the split
This is the most common version, and the easiest to commit by accident. You scale your features, or impute missing values, or select the top features — using the entire dataset — and then split into train and test. The problem: the scaler's mean, the imputer's fill value, the feature selection all now encode information from the test rows. The test set has leaked into the transformations, and your evaluation is optimistic.
# WRONG — the scaler has already seen the test data.
X_scaled = scaler.fit_transform(X_all)
X_train, X_test = split(X_scaled)
# RIGHT — split first; fit the scaler on train only, then apply to test.
X_train, X_test = split(X_all)
scaler.fit(X_train)
X_train = scaler.transform(X_train)
X_test = scaler.transform(X_test) # test never influences the fitThe robust fix is to bundle preprocessing and model into a pipeline and use cross-validation, so every fold re-fits the preprocessing on only that fold's training portion. Then leakage of this kind is structurally impossible, not merely something you remembered to avoid.
Form 2: target leakage
Here a feature secretly contains the answer. Imagine predicting whether a customer will churn, and one of your input columns is cancellation_date — populated only for customers who already churned. The model will "learn" to check that column and score almost perfectly, having discovered nothing except that cancelled customers have cancellation dates. Target leakage usually comes from features created after or because of the outcome you are trying to predict.
Ask of every feature: would I have this at prediction time?
For each input, ask when its value becomes known. If a feature is only populated as a consequence of the outcome — or is filled in after the moment you would actually predict — it is leakage. This one question, asked feature by feature, catches most target leakage before it ever reaches a model.
Form 3: temporal leakage
When your data has a time dimension, using the future to predict the past is a subtle, specific trap. A random train/test split scatters future and past rows on both sides, so the model trains on data from after the events it is tested on — knowledge it could never have in reality. Forecasting problems demand a temporal split: train on earlier data, test on strictly later data, mirroring how the model will actually be used. Duplicate or near-duplicate records straddling the split cause a related problem, quietly placing "the same" example in both train and test.
| Leakage type | How it sneaks in | The fix |
|---|---|---|
| Preprocessing | Transforms fit on the full dataset | Split first; fit on train only; use a pipeline |
| Target | A feature encodes the outcome | Audit features for availability at prediction time |
| Temporal | Random split mixes future into training | Split by time; test on strictly later data |
| Duplicates | Same record in train and test | De-duplicate and group-split related rows |
How to design leakage out
- 1
Split before you touch the data
Make the train/test division the very first step, before any exploration that could influence choices, and before any fitting.
- 2
Fit every learned transform on training only
Scalers, imputers, encoders, feature selectors — all fit on train, then applied to test. A pipeline enforces this across cross-validation folds.
- 3
Interrogate each feature's timing
For every input, confirm its value is known at the moment of a real prediction. Drop or rebuild anything that is not.
- 4
Respect time and identity
Use temporal splits for time-series problems, and group-aware splits so the same entity does not appear on both sides.
Practical takeaway
The instinct to cultivate is suspicion of good news. When a model scores far better than the problem's difficulty suggests it should, the first hypothesis is not "we are brilliant" but "something leaked." Split before you preprocess, fit learned transformations on training data alone, audit every feature for whether you would truly have it at prediction time, and honour the arrow of time when your data has one. Leakage is a bug that flatters you right up until production; the defence is a set of habits that make it structurally hard to commit.
Sources & Further Reading
- 01Leakage and the Reproducibility Crisis in ML-based Science — Kapoor & Narayanan, 2022Documents how leakage has produced over-optimistic, non-reproducible results across many fields.
- 02Common pitfalls and recommended practices — scikit-learnIncludes a direct treatment of data leakage and how to avoid it.
- 03Cross-validation: evaluating estimator performance — scikit-learnHow proper cross-validation keeps preprocessing from leaking across folds.
Editorial note — A conceptual explainer of data leakage and its remedies. Code and the churn/cancellation-date scenario are illustrative; no specific benchmark numbers or study statistics are quoted beyond the cited paper's existence.
