All concepts
Random Forest
Average many decorrelated decision trees to cut variance and boost accuracy.
Classical ML · Intermediate · ~8 min
In plain English
Ask a hundred slightly different experts, each trained on a random sample of the data and allowed to look at only some of the clues, then take the majority vote.
Why it's worth your time
It's the strongest model you can get with almost no tuning, and the honest baseline for any tabular problem.
If you remember three things
- Bagging plus random feature subsets = decorrelated trees
- Averaging many high-variance trees cuts variance, not bias
- Out-of-bag error is free validation
Overview
A random forest trains many decision trees on bootstrap samples, each considering a random subset of features per split. Averaging (or voting) across these decorrelated trees dramatically reduces variance versus a single tree, giving a robust, low-tuning baseline for tabular data.
How it works
- One tree overfits A single deep decision tree fits the training data closely — high variance, unstable.
- Bootstrap samples Each tree is trained on a random resample of the rows, so no two trees see the same data.
- Random features per split At each split a tree only considers a random subset of features — this decorrelates the trees.
- Every tree votes A new point is pushed through all the trees; each casts a prediction.
- Average = low variance Averaging decorrelated trees cancels their individual errors — far more stable than one tree.
In an interview
Random forest is bagging over decision trees with feature subsampling. Each tree sees a bootstrap sample and random features per split, so the trees are decorrelated; averaging their predictions cuts variance and gives a strong, robust tabular baseline.
Production defaults
- Trees
- 300–500. More never hurts accuracy, only latency
- max_features
- sqrt(n) for classification, n/3 for regression — this is what decorrelates the trees
- Depth
- leave unlimited but set min_samples_leaf 1–5; the ensemble handles the variance
What breaks
- Slower than the model it replaced — 500 deep trees is a lot of memory at serve time. Cut tree count and depth — accuracy usually barely moves.
- Worse than a single boosted model — Expected on many tabular tasks. Forests cut variance; boosting cuts bias. Try XGBoost/LightGBM.