All concepts
Feature Scaling
Put features on comparable numeric ranges so distances and gradients behave.
ML Foundations · Beginner · ~8 min
In plain English
One column is measured in rupees and another in years. Without rescaling, the rupees column shouts and the years column whispers — even if years matter more.
Why it's worth your time
It's a one-line fix that silently decides whether KNN, SVM, PCA, K-means and neural nets work at all.
If you remember three things
- Standardize: subtract the mean, divide by the standard deviation
- Fit the scaler on train ONLY, then reuse those numbers everywhere
- Tree models don't care — they split per feature
Overview
Rescales features onto comparable numeric ranges so no single large-magnitude column dominates distances or gradients. Standardization uses z = (x − mean)/std; min-max squeezes to [0,1]. It's essential for distance- and gradient-based methods.
How it works
- Start: Raw Features Income may be in thousands while age is in tens; large-scale columns dominate distance and gradient magnitude.
- Raw Features -> Fit Scaler Compute mean/std or min/max on the training split only.
- Fit Scaler -> Transform Apply the saved scaler to train, validation, and serving data using the same parameters.
- Transform -> Stable Learning KNN, SVM, PCA, logistic regression, and neural nets become easier to optimize and compare.
In an interview
Feature scaling puts columns on comparable ranges, typically z = (x − mean_train)/std_train. Without it, a feature measured in thousands swamps one measured in tens, distorting Euclidean distances and making gradient descent zig-zag. Critically, you fit the scaler on the training split only and reuse those parameters everywhere to avoid leakage.
Production defaults
- Default
- StandardScaler (z-score). Robust scaling when outliers are real and you want to keep them
- Min-max
- only when you need a bounded range and you trust your min/max — one outlier ruins it
- Pipeline it
- scaler inside the CV pipeline, never fit before splitting. This is the #1 leakage source
What breaks
- Model great in notebook, bad in production — The scaler was fit on train+test. Refit inside the split and watch the score drop to the honest number.
- Serving predictions drift over months — Your saved mean/std no longer match reality. Monitor input distributions; rescaling is a retraining trigger.