All concepts
Decision Tree
Recursively split the feature space along the questions that most reduce impurity.
Classical ML · Beginner · ~8 min
In plain English
A flowchart of yes/no questions that a machine wrote for itself, each question chosen to split the remaining examples as cleanly as possible.
Why it's worth your time
It's the only model whose reasoning you can read out loud to a regulator, and it's the building block of the ensembles that still win on tabular data.
If you remember three things
- Each split maximizes purity (Gini or entropy)
- An unconstrained tree memorizes the training set perfectly
- Depth, min-samples-leaf and pruning are the whole tuning surface
Overview
A decision tree carves the feature space into rectangles by asking one threshold question at a time. It greedily picks the split that most reduces impurity (Gini or entropy). Trees are interpretable and non-linear but overfit easily — which is exactly why ensembles exist.
How it works
- All data at the root Start with every example in one node. It's mixed — multiple classes together (high impurity).
- First split Try every feature/threshold and pick the split that most reduces impurity. It becomes the root question.
- Split again Recurse on each child, choosing the next best question. The space is now divided into smaller regions.
- Leaves = regions Stop when a node is pure enough or hits a depth/size limit. Each leaf owns a rectangle of the feature space.
- Predict A new point follows the questions down to a leaf and takes that leaf's majority class (or mean for regression).
In an interview
A decision tree recursively splits the feature space to reduce impurity (Gini or entropy). It's interpretable and non-linear and needs no scaling, but a single tree overfits and is high-variance, so we use ensembles like random forests or gradient boosting.
Production defaults
- Depth
- 3–8 for a tree you intend to read; deeper only inside an ensemble
- min_samples_leaf
- ≥ 20 on real data — leaves of 1 are memorized noise
- Imbalance
- class_weight='balanced' before you touch the sampling
What breaks
- 100% train accuracy, mediocre test — The classic single-tree failure. Constrain depth/leaf size, or use a forest — a single tree is a high-variance model by nature.
- Feature importances look wrong — Impurity importance is biased toward high-cardinality features. Use permutation importance instead.