All concepts
Naive Bayes
Use Bayes rule with a simplifying independence assumption for fast classification.
Classical ML · Beginner · ~8 min
In plain English
Judge an email by counting which words show up in spam versus normal mail, pretending each word is independent of the others. It's a lie, and it works anyway.
Why it's worth your time
It trains in one pass over the data and is still a genuinely hard baseline to beat on text classification.
If you remember three things
- Combines a prior with per-feature likelihoods
- Work in log space — products of probabilities underflow
- Only needs the right argmax, which is why the false independence assumption survives
Overview
A probabilistic classifier applying Bayes' rule with the 'naive' assumption that features are conditionally independent given the class. This makes scoring a fast sum of log-probabilities and needs little data, giving a strong, cheap text-classification baseline.
How it works
- Start: Document A text example is converted into token counts or binary word features.
- Document -> Class Priors Estimate how common each class is from the training data.
- Class Priors -> Word Likelihoods Estimate P(word | class), usually with smoothing so unseen words do not zero out the class.
- Word Likelihoods -> Posterior Multiply priors and likelihoods in log-space to score every class.
- Posterior -> Best Class Despite the naive assumption, it is a very strong baseline for spam, intent, and topic classification.
In an interview
Naive Bayes picks the class maximizing P(class)·∏P(feature|class), assuming features are independent given the class. In log-space that's argmax_c [log P(c) + Σ log P(xᵢ|c)]. The independence assumption is usually false, but the classifier is fast, needs little data, and is a very strong baseline for spam and topic classification.
Production defaults
- Smoothing
- Laplace α=1. Without it, one unseen word zeroes an entire class
- Variant
- multinomial for counts, Bernoulli for presence/absence, Gaussian for continuous features
- Use it as
- a baseline you must beat before shipping anything heavier for text
What breaks
- Probabilities are wildly overconfident — Expected — correlated features get double-counted. Calibrate if you need real probabilities, or just use the ranking.
- One class always wins — Severe prior imbalance, or missing smoothing. Check class priors first.