All concepts
Training-Serving Skew
A mismatch between training-time data processing and production-time serving.
MLOps & LLMOps · Intermediate · ~8 min
In plain English
The features you compute at training time and the ones you compute when serving are subtly different — so the model sees a world it never learned.
Why it's worth your time
It's the classic 'great in the notebook, bad in production' bug, and it's usually invisible until you go looking.
If you remember three things
- Different code paths for the same feature is the root cause
- Also caused by leakage: a feature not available at prediction time
- Logging serving features is how you find it
Overview
Training-serving skew is a mismatch between how features are computed offline for training and how they're computed online at serving. Even tiny differences — a timezone, a null default, a scaling constant — silently degrade production quality because the model sees inputs subtly different from what it learned on.
How it works
- Start: Training Pipeline Features are computed one way during offline training.
- Training Pipeline -> Serving Pipeline Production code may compute them differently or with different freshness.
- Serving Pipeline -> Mismatch Small logic, timezone, null, or scaling differences can silently break quality.
- Mismatch -> Skew Tests Compare offline and online feature values for the same entity/time.
- Skew Tests -> Shared Logic Feature stores, contracts, and CI checks reduce skew.
In an interview
Skew is when the same feature is computed differently in the training pipeline than in the serving pipeline — different code, freshness, timezone, null handling, or normalization. The model was fit on one version and scores the other, so quality drops with no error. The fix is shared feature logic (a feature store), schema contracts, and CI checks that compare offline vs online values.
Production defaults
- Share code
- one implementation used by both paths. A feature store enforces this structurally
- Log and compare
- sample serving features and diff their distributions against training
- Availability check
- for every feature, ask 'is this known at prediction time?' before it goes in
What breaks
- Offline 0.92, online 0.71 — Log serving features for a sample and compare directly against training. The gap is usually one or two columns.
- Fine at launch, degrading over time — That's drift, not skew. Different diagnosis — compare against the training window, not against serving.