All concepts
Gradient Boosting / XGBoost
Add shallow trees one at a time, each correcting the residual errors of the last.
Classical ML · Advanced · ~8 min
In plain English
Instead of a hundred independent experts, you train them in a line: each new one studies only the mistakes the group has made so far.
Why it's worth your time
It is still the model that wins tabular competitions and quietly runs most production ranking and risk systems.
If you remember three things
- Each tree fits the residual errors of the ensemble so far
- Learning rate and tree count trade off directly against each other
- Built-in regularization is why it beats plain gradient boosting
Overview
Gradient boosting builds an additive model of shallow trees, where each new tree fits the negative gradient (residuals) of the loss so far. XGBoost/LightGBM add regularization, second-order gradients, and system optimizations, making them the go-to for tabular competitions and production.
How it works
- Start with a weak guess Begin with a simple prediction (e.g. the mean). It's wrong almost everywhere.
- Look at the residuals The residuals are how far off we are at each point — the errors left to fix.
- Fit a small tree to the errors A shallow tree is trained to predict those residuals — it learns where and how we're wrong.
- Add it (shrunk by the learning rate) The new tree's correction is added to the model, scaled down by the learning rate for stability.
- Repeat → a strong model Each new tree corrects the previous ensemble's mistakes. Bias falls step by step — this is boosting.
In an interview
Gradient boosting fits shallow trees sequentially, each correcting the previous ensemble's residuals by following the loss gradient. XGBoost adds regularization and second-order info. It reduces bias and typically tops tabular benchmarks, but needs careful tuning to avoid overfitting.
Production defaults
- LR / trees
- learning_rate 0.05 with early stopping on a validation set — let it pick n_estimators
- Depth
- max_depth 4–8. Deeper overfits fast on tabular data
- Sampling
- subsample 0.8, colsample_bytree 0.8 — cheap variance reduction
- Regularization
- min_child_weight 1–10, reg_lambda 1. Raise both when validation diverges
What breaks
- Train loss keeps falling, validation rises — Textbook boosting overfit. Early stopping is not optional — set it before tuning anything else.
- Great offline, poor online — Boosting will happily exploit a leaky feature. Audit features for anything computed after the label.