All concepts
BM25 / TF-IDF
Score documents by lexical term overlap with saturation and rarity weighting.
RAG & Retrieval · Beginner · ~8 min
In plain English
Score a document by how often the query's words appear in it, discounting words that appear everywhere, and stopping long documents from winning just by being long.
Why it's worth your time
It's a 30-year-old algorithm that still beats embeddings on exact terms — and it's half of every serious retrieval system.
If you remember three things
- TF: term appears often here. IDF: term is rare overall
- BM25 adds saturation and length normalization on top of TF-IDF
- No training, no embeddings, no GPU
Overview
Classic lexical ranking. Documents score by term overlap with the query, weighting rare terms higher (IDF) and saturating repeated terms so frequency stops helping past a point. Unbeatable for exact names, codes, and error strings; often fused with vector search in hybrid RAG.
How it works
- Start: Query Terms Keyword search starts from exact terms in the query.
- Query Terms -> Term Rarity Rare terms receive higher weight because they are more informative.
- Term Rarity -> Term Frequency BM25 saturates term frequency so repeated words do not dominate forever.
- Term Frequency -> Keyword Rank Lexical ranking is excellent for exact names, codes, error messages, and rare terms.
- Keyword Rank -> Hybrid Fuse Production RAG often combines BM25 with vector retrieval for recall and precision.
In an interview
BM25 is a bag-of-words ranking function: score(d,q)=Σ IDF(t)·BM25_tf(t,d). It rewards rare query terms via IDF and applies saturating term frequency with length normalization, so a word repeated ten times isn't ten times better. It's the sparse-retrieval baseline — great for exact tokens, weak on synonyms and paraphrase.
Production defaults
- Parameters
- k1 = 1.2, b = 0.75. These defaults are good; tune only with a measured reason
- Analyzer
- the part that actually matters. Get stemming, casing and identifier handling right before touching k1/b
- Always pair
- with vector search. Neither is sufficient alone on a real corpus
What breaks
- Paraphrased queries find nothing — Expected — lexical matching only. That's exactly what the vector half of hybrid retrieval is for.
- Long documents dominate — b too low. Raise toward 0.75+ so length normalization actually applies.