Skip to content

Data Leakage in Machine Learning

The most common reason models look brilliant in testing and fail in production — and how to prevent it.

Editorial team 2 min read

Data leakage happens when information that won't be available at prediction time sneaks into training or evaluation. The model looks excellent in testing, then disappoints in production.

Common Forms

  • Target leakage: a feature that is a consequence of the outcome. Predicting loan default using "number of collection calls" leaks the answer.
  • Train–test contamination: preprocessing (scaling, imputation, feature selection) fitted on the full dataset before splitting.
  • Temporal leakage: random splits of time-ordered data let the model learn from the future.
  • Group leakage: records from the same person or device on both sides of the split.
  • Duplicate leakage: near-identical rows in training and test sets.

Warning Signs

  • Performance that seems too good to be true.
  • One feature with overwhelming importance.
  • A big drop between offline scores and live results.

Prevention Checklist

  1. For every feature, ask: would I know this at the moment of prediction?
  2. Split data before any preprocessing, and fit transformations on training data only — use pipelines.
  3. Split by time for temporal problems and by group for grouped data.
  4. Remove duplicates across splits.
  5. Keep the test set locked away until the end.

A Useful Habit

Write down the exact moment a prediction will be made in production, and the data available at that moment. Build the training dataset to mirror it.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026