All concepts
Cross-Validation
Rotate which fold is held out for validation, then average — a stable estimate from limited data.
ML Foundations · Beginner · ~8 min
In plain English
Instead of one exam, you sit five — each time a different fifth of the questions is held back. Your real score is the average, and you learn how much it wobbles.
Why it's worth your time
A single train/test split on a small dataset can be lucky by several points. Cross-validation is how you stop fooling yourself.
If you remember three things
- k=5 or 10 is the standard; higher k costs compute for little gain
- Stratify for classification so every fold has the same class mix
- The spread across folds matters as much as the mean
Overview
A single train/validation split is noisy and wastes data. k-fold cross-validation splits the data into k folds, trains on k−1 and validates on the held-out fold, rotates through all folds, and averages the scores — a lower-variance estimate of generalization.
How it works
- One dataset We have limited data and need an honest estimate of real-world performance.
- Split into k folds Divide the data into k equal, disjoint folds (commonly k=5 or 10).
- Round 1 Hold out fold 1 to validate; train on the other k−1 folds; record the score.
- Rotate Repeat so every fold serves as validation exactly once — every point is used for both training and validation.
- Average Average the k scores for a stable, low-variance estimate (and a spread to gauge stability).
In an interview
k-fold cross-validation splits data into k folds, trains on k−1 and validates on the held-out one, rotates through all folds, and averages the scores. It gives a lower-variance estimate of generalization than a single split and uses all data for both training and validation — crucial when data is limited.
Production defaults
- Folds
- 5 for most work, 10 when data is scarce and compute is cheap
- Time series
- never shuffle. Use forward-chaining splits, train on the past, test on the future
- Grouped data
- split by group (user, patient, device) or the same entity leaks across folds
What breaks
- CV score far better than production — Preprocessing was fit on all the data. Every transform must live inside the fold, in a pipeline.
- Folds disagree wildly — Too little data or a genuinely heterogeneous population. Report the spread; don't hide it behind a mean.