All concepts

Random Forest

Average many decorrelated decision trees to cut variance and boost accuracy.

Classical ML · Intermediate · ~8 min

In plain English

Ask a hundred slightly different experts, each trained on a random sample of the data and allowed to look at only some of the clues, then take the majority vote.

Why it's worth your time

It's the strongest model you can get with almost no tuning, and the honest baseline for any tabular problem.

If you remember three things

  • Bagging plus random feature subsets = decorrelated trees
  • Averaging many high-variance trees cuts variance, not bias
  • Out-of-bag error is free validation

Overview

A random forest trains many decision trees on bootstrap samples, each considering a random subset of features per split. Averaging (or voting) across these decorrelated trees dramatically reduces variance versus a single tree, giving a robust, low-tuning baseline for tabular data.

How it works

  1. One tree overfits A single deep decision tree fits the training data closely — high variance, unstable.
  2. Bootstrap samples Each tree is trained on a random resample of the rows, so no two trees see the same data.
  3. Random features per split At each split a tree only considers a random subset of features — this decorrelates the trees.
  4. Every tree votes A new point is pushed through all the trees; each casts a prediction.
  5. Average = low variance Averaging decorrelated trees cancels their individual errors — far more stable than one tree.

In an interview

Random forest is bagging over decision trees with feature subsampling. Each tree sees a bootstrap sample and random features per split, so the trees are decorrelated; averaging their predictions cuts variance and gives a strong, robust tabular baseline.

Production defaults

Trees
300–500. More never hurts accuracy, only latency
max_features
sqrt(n) for classification, n/3 for regression — this is what decorrelates the trees
Depth
leave unlimited but set min_samples_leaf 1–5; the ensemble handles the variance

What breaks

  • Slower than the model it replaced — 500 deep trees is a lot of memory at serve time. Cut tree count and depth — accuracy usually barely moves.
  • Worse than a single boosted model — Expected on many tabular tasks. Forests cut variance; boosting cuts bias. Try XGBoost/LightGBM.

Watch it explained

What is Random Forest? — IBM Technology, 5:21

Related