All concepts

BM25 / TF-IDF

Score documents by lexical term overlap with saturation and rarity weighting.

RAG & Retrieval · Beginner · ~8 min

In plain English

Score a document by how often the query's words appear in it, discounting words that appear everywhere, and stopping long documents from winning just by being long.

Why it's worth your time

It's a 30-year-old algorithm that still beats embeddings on exact terms — and it's half of every serious retrieval system.

If you remember three things

  • TF: term appears often here. IDF: term is rare overall
  • BM25 adds saturation and length normalization on top of TF-IDF
  • No training, no embeddings, no GPU

Overview

Classic lexical ranking. Documents score by term overlap with the query, weighting rare terms higher (IDF) and saturating repeated terms so frequency stops helping past a point. Unbeatable for exact names, codes, and error strings; often fused with vector search in hybrid RAG.

How it works

  1. Start: Query Terms Keyword search starts from exact terms in the query.
  2. Query Terms -> Term Rarity Rare terms receive higher weight because they are more informative.
  3. Term Rarity -> Term Frequency BM25 saturates term frequency so repeated words do not dominate forever.
  4. Term Frequency -> Keyword Rank Lexical ranking is excellent for exact names, codes, error messages, and rare terms.
  5. Keyword Rank -> Hybrid Fuse Production RAG often combines BM25 with vector retrieval for recall and precision.

In an interview

BM25 is a bag-of-words ranking function: score(d,q)=Σ IDF(t)·BM25_tf(t,d). It rewards rare query terms via IDF and applies saturating term frequency with length normalization, so a word repeated ten times isn't ten times better. It's the sparse-retrieval baseline — great for exact tokens, weak on synonyms and paraphrase.

Production defaults

Parameters
k1 = 1.2, b = 0.75. These defaults are good; tune only with a measured reason
Analyzer
the part that actually matters. Get stemming, casing and identifier handling right before touching k1/b
Always pair
with vector search. Neither is sufficient alone on a real corpus

What breaks

  • Paraphrased queries find nothing — Expected — lexical matching only. That's exactly what the vector half of hybrid retrieval is for.
  • Long documents dominate — b too low. Raise toward 0.75+ so length normalization actually applies.

Watch it explained

Natural Language Processing|TF-IDF Intuition| Text Prerocessing — Krish Naik, 8:27

Related