All concepts

Naive Bayes

Use Bayes rule with a simplifying independence assumption for fast classification.

Classical ML · Beginner · ~8 min

In plain English

Judge an email by counting which words show up in spam versus normal mail, pretending each word is independent of the others. It's a lie, and it works anyway.

Why it's worth your time

It trains in one pass over the data and is still a genuinely hard baseline to beat on text classification.

If you remember three things

  • Combines a prior with per-feature likelihoods
  • Work in log space — products of probabilities underflow
  • Only needs the right argmax, which is why the false independence assumption survives

Overview

A probabilistic classifier applying Bayes' rule with the 'naive' assumption that features are conditionally independent given the class. This makes scoring a fast sum of log-probabilities and needs little data, giving a strong, cheap text-classification baseline.

How it works

  1. Start: Document A text example is converted into token counts or binary word features.
  2. Document -> Class Priors Estimate how common each class is from the training data.
  3. Class Priors -> Word Likelihoods Estimate P(word | class), usually with smoothing so unseen words do not zero out the class.
  4. Word Likelihoods -> Posterior Multiply priors and likelihoods in log-space to score every class.
  5. Posterior -> Best Class Despite the naive assumption, it is a very strong baseline for spam, intent, and topic classification.

In an interview

Naive Bayes picks the class maximizing P(class)·∏P(feature|class), assuming features are independent given the class. In log-space that's argmax_c [log P(c) + Σ log P(xᵢ|c)]. The independence assumption is usually false, but the classifier is fast, needs little data, and is a very strong baseline for spam and topic classification.

Production defaults

Smoothing
Laplace α=1. Without it, one unseen word zeroes an entire class
Variant
multinomial for counts, Bernoulli for presence/absence, Gaussian for continuous features
Use it as
a baseline you must beat before shipping anything heavier for text

What breaks

  • Probabilities are wildly overconfident — Expected — correlated features get double-counted. Calibrate if you need real probabilities, or just use the ranking.
  • One class always wins — Severe prior imbalance, or missing smoothing. Check class priors first.

Watch it explained

Naive bayes classifier explained for beginners — AI Explored, 5:02

Related