All concepts

Data Drift

Detect when production input distributions move away from training data.

MLOps & LLMOps · Intermediate · ~8 min

In plain English

The world moves on but your model doesn't. The inputs it sees today no longer look like the ones it learned from.

Why it's worth your time

It's the most common reason a model that worked at launch quietly stops working, months later, with no error anywhere.

If you remember three things

  • Covariate drift: inputs changed. Concept drift: the relationship changed
  • Drift is detectable long before labels arrive
  • Detection is cheap; retraining is not — set a threshold

Overview

Data drift is when the distribution of production inputs moves away from the training distribution, so a model scores fresh data the loss surface never covered. It degrades quality silently — no error is thrown, just worse predictions — which is why input distributions are monitored separately from accuracy.

How it works

  1. Start: Training Data The model learned patterns from a historical distribution.
  2. Training Data -> Production Data Live traffic may change because users, seasonality, or upstream systems change.
  3. Production Data -> Distribution Check Monitor feature histograms, embeddings, PSI, KS tests, and missing rates.
  4. Distribution Check -> Alert Significant drift triggers investigation.
  5. Alert -> Retrain / Fix Decide whether to retrain, adjust features, or fix broken pipelines.

In an interview

Data drift is a shift in the input distribution P(x) between training and serving. You detect it without labels by watching feature histograms and embeddings with tests like PSI (>0.2 is a red flag) or the KS statistic. It matters because it silently erodes accuracy long before labeled outcomes arrive to confirm the drop.

Production defaults

Monitor
input feature distributions (PSI or KS test) plus prediction distribution
Threshold
PSI > 0.2 on a key feature is a common investigate signal
Label lag
when labels arrive late, drift on inputs is your only early warning
Response
investigate first. Not all drift hurts accuracy, and retraining on drifted data can make it worse

What breaks

  • Accuracy fell with no code change — Compare input distributions against the training window before touching the model.
  • Constant drift alerts nobody acts on — Thresholds too tight, or you're alerting on features that don't matter. Weight by feature importance.

Watch it explained

Model Drift In machine learning | Concept drift vs Data Drift in Machine Learning — Unfold Data Science, 8:42

Related