All concepts
One-Hot Encoding
Represent a category as a sparse vector with exactly one active position.
ML Foundations · Beginner · ~8 min
In plain English
Instead of numbering categories 1, 2, 3 — which secretly tells the model 3 is bigger than 1 — you give each category its own on/off switch.
Why it's worth your time
Integer-encoding a category is one of the most common silent bugs in tabular ML: the model learns an ordering you never meant to imply.
If you remember three things
- One column per category, exactly one of them set to 1
- Fix the vocabulary on the training split — it's a serving contract
- High cardinality explodes the width; that's the whole trade-off
Overview
Encodes a categorical value as a vector that is all zeros except a single 1 at the category's index. It gives linear and tree models a numeric input without imposing a fake order between categories, unlike integer label encoding.
How it works
- Start: Category Start with a symbolic category such as plan=premium; the model cannot multiply raw strings.
- Category -> Vocabulary Create a fixed list of allowed categories. The index is the contract shared by training and serving.
- Vocabulary -> Sparse Vector Turn the selected category into [0,0,1,0]. Only one bit is on.
- Sparse Vector -> Model Input Concatenate the one-hot vector with numeric features so linear and tree models can consume it.
In an interview
One-hot encoding maps each category to its own binary column, setting one position to 1 and the rest to 0. It avoids the false ordinal relationship that integer encoding creates, at the cost of one dimension per category, so high-cardinality features blow up the input width.
Production defaults
- Cardinality limit
- one-hot up to ~50 categories; above that use target encoding, hashing, or learned embeddings
- Unknown bucket
- always reserve one. A category unseen in training will appear in production on day one
- Trees
- tree models handle categorical splits natively in LightGBM/CatBoost — one-hot can actually hurt them
What breaks
- Serving crashes on an unseen category — You had no unknown bucket. Add it and set handle_unknown='ignore' equivalents.
- Column order differs between train and serve — Persist the fitted encoder, not the column list you wrote down.