012 min
The problem data leakage creates
The name collides with a security term, so start by separating them. A data breach is confidential information escaping to people who should not have it. Data leakage in machine learning runs the other way: information arriving into a model that should not have been available to it. Nothing escapes and nobody is attacked. This entry is about the second meaning throughout.
The reason the concept needs a name is that the standard safeguard does not catch it.
Every machine learning workflow holds back a test set, trains on the rest, and reports performance on the held-out data. That procedure is designed to answer one question: will this model work on data it has not seen? Leakage breaks the procedure at its foundation, because the test set is no longer data the model has not seen. Some trace of the answer was already in the training data.
What makes this the most expensive failure in applied machine learning is the shape of the feedback. A model with a bug crashes. A model that underfits scores badly and gets fixed. A model with leakage scores better than the honest version, which means every incentive in the process points toward shipping it. The number on the slide goes up. The review goes well. The failure surfaces in production, weeks later, where the diagnosis is hardest.
Shachar Kaufman, Saharon Rosset, Claudia Perlich and Ori Stitelman gave the canonical treatment in ACM Transactions on Knowledge Discovery from Data in 2012, defining leakage as the introduction of information about the target which should not legitimately be available to mine from. Their motivation was that the problem kept appearing in major public competitions and in real projects, while the literature had left it largely unexamined.
022 min
How data leakage works
Follow one running example: a model that predicts which customers will cancel a subscription in the next 30 days.
The training data is one row per customer, with features drawn from their account, and a label saying whether they cancelled. The model reaches 0.95 AUC on the held-out test set, which is far better than anyone expected.
Where the information came from
One of the features is support_tickets_last_90d. It was computed when the dataset was assembled, in March, from the current state of the support system. The labels cover cancellations through February.
So for a customer who cancelled in February, that feature counts tickets they filed while cancelling β the complaint, the refund request, the downgrade question. The feature contains the answer. At prediction time the model will be asked about customers who have not cancelled yet, and those tickets will not exist.
The model learned a rule that is genuinely true in the training data and genuinely useless in production: customers with a sudden burst of support tickets are about to cancel, because in this dataset the burst is the cancellation.
The general shape
Every instance of leakage has the same structure. There is a piece of information, it is correlated with the target, and it would not be available at the moment the prediction has to be made. The reason it is hard to catch is that only the third condition is violated, and the third condition is about time and process rather than about the data.
A correlation matrix will not flag it, because the correlation is real. Cross-validation will not flag it, because every fold contains the same contamination. A held-out test set will not flag it, because the leak is in the features rather than in the split.
The subtler route
Leakage also arrives through preprocessing, and this version catches careful people. Suppose you scale features by computing the mean and standard deviation across the whole dataset before splitting. Those statistics contain information from the test rows. The effect is usually small, but it is the same failure: the training process saw something the deployed model will not have.
The fix Kaufman and colleagues propose is structural rather than a checklist. They call it a learn-predict separation: build the dataset so that everything available to the model is explicitly tied to a point in time before the prediction would be made.
032 min
A concrete example: the patient identifier
The case Kaufman and colleagues use is worth knowing because nobody could reasonably have predicted it.
KDD Cup 2008 was a competition on detecting cancer from mammography data. Analysing the data afterwards, the authors point out that the patient identifier β a field most competitors ignored, as most people ignore identifiers β carried tremendous and unexpected predictive power.
The explanation is mundane and that is the point. The dataset had been assembled from multiple sources: different clinical studies, institutions or equipment. Patient identifiers were assigned in consecutive blocks per source. Because the sources differed in how many cancer cases they contained, the identifier range encoded which source a record came from, and therefore carried information about the outcome. As the paper puts it, the merge was done without obfuscating the source.
No clinical information leaked. No rule was broken by any competitor. A model using that feature would have scored well on the competition's test set and been worthless on a new hospital's data, where identifiers mean something else entirely.
The same paper describes a second competition the same year, the INFORMS Data Mining Challenge 2008 on pneumonia diagnosis from hospital records. There the target had originally been embedded as a special value of one or more features. The organisers removed those values before release β and the removal itself was detectable. A record with all condition codes missing was recognisably a record that had been cleaned, which made it recognisably a positive case.
Two different mechanisms, one lesson. Leakage attaches to how the dataset was built, not to what the fields are supposed to mean.
042 min
The forms data leakage takes
Kapoor and Narayanan's 2022 survey proposes a taxonomy of eight leakage types, ranging from textbook errors to open research problems. The table below names the forms that account for most real incidents, with the question that detects each.
| Form | What happens | The question that catches it |
|---|---|---|
| Target leakage | A feature contains information created by or after the outcome | For each feature, when was this value actually recorded? |
| Temporal leakage | Training data includes events that happened after the test period | Is the split by time, or by random row? |
| Group leakage | Rows from the same entity appear in both training and test | Is the split by patient, user or session, or by row? |
| Preprocessing leakage | Scaling, imputation or feature selection computed before the split | Was this statistic computed on training rows only? |
| Duplicate leakage | The same record appears in both sets, sometimes in altered form | Have you checked for near-duplicates, not just exact ones? |
| Benchmark contamination | Test items appeared in a model's pretraining corpus | Could this evaluation set have been on the public web? |
Group leakage is the one that most often survives review, because the split looks correct. A medical dataset with several scans per patient, split randomly by scan, puts the same patient on both sides. The model can recognise the patient rather than the disease, and the test score is measuring memory.
052 min
A second case: an entire field at once
The patient-identifier case is one dataset. The COVID-19 imaging literature shows what happens when the same failure runs through a whole field under time pressure.
Michael Roberts and colleagues published a systematic review in Nature Machine Intelligence in 2021, covering machine learning for detecting and prognosticating COVID-19 from chest radiographs and CT scans. They identified 2,212 studies, screened 415 in, and reviewed 62 in detail. None of the 62 was of potential clinical use, because of methodological flaws.
The specific flaws read like the taxonomy above made real. Some studies used images of children as the non-COVID class and images of adults as the COVID class, so, as Roberts put it, all the model could usefully do was tell the difference between children and adults. Public datasets had been merged and re-merged over time into what the review calls Frankenstein datasets, so that reproduction was impossible and the same images could appear on both sides of a split. Models were trained and tested on the same data. Datasets frequently came from a single hospital with no geographic variation.
Alex DeGrave, Joseph Janizek and Su-In Lee examined the same class of systems in the same journal in the same year and found the mechanism directly. Using explainable-AI techniques, they showed the models were relying on confounding factors rather than pathology: text markers and patient positioning specific to each dataset. Their summary is the one to remember β the systems appear accurate, and fail when tested in new hospitals.
One variable separates this from the KDD Cup case: whether anyone was checking. The competition dataset was scrutinised by hundreds of competitors and the leak was found and written up. The COVID literature was produced fast, reviewed fast, and published across hundreds of venues, and the leakage was found years later by people doing a systematic review.
What the contrast teaches is not available from either case alone. Leakage is not primarily a skill problem. It is a review problem, and it survives in exact proportion to how little anyone is incentivised to look for it.
062 min
What data leakage means for your work
For an engineer: split by entity and by time, never by row. If your data has users, patients, sessions or devices, the split goes on that key. If predictions will be made about the future, the split goes on a date. Random row splits are the default in every tutorial and are wrong for most production problems.
For an engineer, second: audit features by timestamp, not by name. For every feature, ask when the value was written, not what it represents. account_status sounds harmless until you learn it is updated on cancellation. This audit is tedious and it is the single highest-yield hour in a modelling project.
For a product manager: treat a surprisingly good result as a bug report. The strongest practical signal for leakage is a model that substantially beats what domain experts thought possible. That reaction is usually celebrated. It should trigger an investigation, and the investigation should happen before the result is presented to anyone else.
For a founder: a benchmark number is a claim about a dataset, not about a product. When evaluating a vendor or an acquisition, ask how the evaluation set was constructed and split, and whether the model was ever exposed to it. The answer is diagnostic, including when there is no answer.
For anyone shipping on top of a general model: benchmark contamination is the version of this that reaches you. A model that has seen a public benchmark during pretraining will report a score that overstates its ability on genuinely new problems of the same type. Build a small private evaluation set from your own data, keep it off the public web, and treat it as the only number that describes your use case.
072 min
How to detect and reduce data leakage
Leakage is found by procedure, not by intuition, and every method below is cheap relative to discovering it in production.
Start from the prediction moment. Fix the exact instant the model will be called in production. Then require of every feature that its value was knowable at that instant. This is the learn-predict separation Kaufman and colleagues recommend, and it turns a vague worry into a per-feature question with a yes or no answer.
Use exploratory analysis as a detector. The same paper recommends this, and the logic is specific: leakage shows up as a pattern in the data that is surprising. An identifier that correlates with the target is surprising. A feature that is almost perfectly predictive on its own is surprising. Surprise is not proof, since legitimate strong predictors exist, but there are few enough surprises in any dataset that checking each one is affordable.
Rank features by importance and interrogate the top of the list. If the single most important feature is one you would not have expected, that is where to look first.
Hold out a set the pipeline never touches. Separate from cross-validation, keep data from a later time period, a different site, or a different cohort, and score against it once, late. A gap between that score and your validation score is the size of your leakage plus your distribution shift, and it is better to learn it there than in production.
Write down the provenance. Kapoor and Narayanan propose model info sheets for reporting what a model was trained and evaluated on, precisely because none of the errors they found could have been caught by reading the papers. The internal version of this is a short document per model recording where each feature came from, when it is written, and how the split was made.
Treat it as a review responsibility. The COVID review above is the argument. Give someone other than the model's author the explicit job of looking for leakage, and give them the provenance document rather than the notebook.
082 min
Data leakage vs. nearby concepts
| Compared with | The one fact that decides |
|---|---|
| Overfitting | Which score is inflated. An overfit model scores well on training data and badly on held-out data, so the test set catches it. Leakage inflates the held-out score too, which is why the standard safeguard does not fire. |
| Shortcut learning | Cause and effect. Shortcut learning is a model relying on an easy correlate rather than the intended signal. Leakage is one way that correlate gets there. The COVID models showed both at once: leakage put the dataset markers in, shortcut learning is what the model then did with them. |
| Distribution shift | Whether the training data was ever honest. Under distribution shift the training data was a fair sample and the world changed afterwards. Under leakage the training data was never a fair sample of what prediction time looks like. One is a maintenance problem; the other is a construction defect. |
| Data breach | Different fields entirely. A breach is confidential data escaping to unauthorised people, a security failure. Data leakage is information entering a model that should not have been available to it. They share a word and nothing else. |
The question that separates leakage from overfitting fastest: does the model also score well on the held-out set? If yes, and production disagrees, stop tuning regularisation and start auditing where the features came from.
092 min
Where the evidence is contested
The disagreement worth reporting is about scale rather than existence, and it has moved recently.
Sayash Kapoor and Arvind Narayanan make the strongest claim. Surveying research communities that adopted machine learning methods, they found errors in 17 fields, collectively affecting 329 papers, and in some cases leading to conclusions they describe as wildly overoptimistic. They then ran a reproducibility study in civil war prediction, a field where complex models were believed to substantially beat older statistical methods. Every paper claiming that superiority failed to reproduce once leakage was accounted for, and the complex models did not perform substantively better than decades-old logistic regression.
The objection to a survey of this kind is selection: it counts fields where somebody went looking and found something, which says less about the base rate than the headline suggests. The authors are explicit that their count comes from a survey of known errors, and a fair reading treats 329 as a floor rather than an estimate.
The live version of this argument has moved to language models, where leakage takes the form of benchmark contamination. Hugh Zhang and colleagues at Scale AI built GSM1k in 2024, a fresh benchmark matched to the established GSM8k on human solve rates, solution steps and answer magnitude. They report accuracy drops of up to 8 percent on the new benchmark, with several model families showing evidence of systematic overfitting across almost all sizes, and a positive correlation between a model's probability of generating GSM8k examples and its performance gap.
Importantly for anyone tempted to use this as a blanket dismissal: the same paper reports that frontier models showed minimal signs of overfitting and demonstrated generalisation to novel problems. The honest summary is that contamination is real, measurable and uneven, rather than that all benchmark numbers are fiction.
Where this leaves a practitioner is unchanged by the outcome. Whatever the true base rate, the cost of checking your own pipeline is a few hours, and the cost of not checking is discovered by users.
?8 questions
Questions people ask
What is data leakage in machine learning?
Is data leakage the same as a data breach?
How is data leakage different from overfitting?
What is a real example of data leakage?
How do you detect data leakage?
How common is data leakage?
What is benchmark contamination?
Why does splitting data randomly cause leakage?
Β§6 sources
Sources
Kaufman, S., Rosset, S., Perlich, C. and Stitelman, O. (2012). Leakage in Data Mining: Formulation, Detection, and Avoidance. ACM Transactions on Knowledge Discovery from Data 6(4), Article 15. DOI 10.1145/2382577.2382579
Kapoor, S. and Narayanan, A. (2023). Leakage and the Reproducibility Crisis in ML-based Science. Patterns 4(9), 100804. arXiv:2207.07048
Roberts, M. et al. (2021). Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nature Machine Intelligence 3, 199-217
DeGrave, A. J., Janizek, J. D. and Lee, S.-I. (2021). AI for radiographic COVID-19 detection selects shortcuts over signal. Nature Machine Intelligence 3, 610-619
Show all 6 sourcesShow fewer sources
Zhang, H. et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332
University of Cambridge (2021). Machine learning models for diagnosing COVID-19 are not yet suitable for clinical use





