All concepts

One-Hot Encoding

Represent a category as a sparse vector with exactly one active position.

ML Foundations · Beginner · ~8 min

In plain English

Instead of numbering categories 1, 2, 3 — which secretly tells the model 3 is bigger than 1 — you give each category its own on/off switch.

Why it's worth your time

Integer-encoding a category is one of the most common silent bugs in tabular ML: the model learns an ordering you never meant to imply.

If you remember three things

  • One column per category, exactly one of them set to 1
  • Fix the vocabulary on the training split — it's a serving contract
  • High cardinality explodes the width; that's the whole trade-off

Overview

Encodes a categorical value as a vector that is all zeros except a single 1 at the category's index. It gives linear and tree models a numeric input without imposing a fake order between categories, unlike integer label encoding.

How it works

  1. Start: Category Start with a symbolic category such as plan=premium; the model cannot multiply raw strings.
  2. Category -> Vocabulary Create a fixed list of allowed categories. The index is the contract shared by training and serving.
  3. Vocabulary -> Sparse Vector Turn the selected category into [0,0,1,0]. Only one bit is on.
  4. Sparse Vector -> Model Input Concatenate the one-hot vector with numeric features so linear and tree models can consume it.

In an interview

One-hot encoding maps each category to its own binary column, setting one position to 1 and the rest to 0. It avoids the false ordinal relationship that integer encoding creates, at the cost of one dimension per category, so high-cardinality features blow up the input width.

Production defaults

Cardinality limit
one-hot up to ~50 categories; above that use target encoding, hashing, or learned embeddings
Unknown bucket
always reserve one. A category unseen in training will appear in production on day one
Trees
tree models handle categorical splits natively in LightGBM/CatBoost — one-hot can actually hurt them

What breaks

  • Serving crashes on an unseen category — You had no unknown bucket. Add it and set handle_unknown='ignore' equivalents.
  • Column order differs between train and serve — Persist the fitted encoder, not the column list you wrote down.

Watch it explained

Principles behind neural networks and one hot encoding — Google for Developers, 7:31

Related