All concepts
Multistage Ranking
Rank items in stages: broad retrieval, lightweight scoring, heavy reranking, and final business rules.
Applied ML · Advanced · ~8 min
In plain English
Millions of candidates, one screen of results. Narrow it in stages: something crude and instant first, something expensive and accurate at the end.
Why it's worth your time
It's the architecture of every search, feed and recommendation system — and the same shape as retrieve-then-rerank in RAG.
If you remember three things
- Retrieval → ranking → re-ranking, each smaller and more expensive
- Recall is the earlier stage's job; precision is the later stage's
- A candidate lost early can never be recovered
Overview
A funnel that ranks huge candidate sets in stages: cheap retrieval narrows millions to thousands, a stronger ranker scores those with richer features, and a precise reranker applies diversity, freshness, and policy. Each stage trades recall for precision so heavy models only touch a small set.
How it works
- Start: Huge Corpus Millions of candidates cannot all go through an expensive model.
- Huge Corpus -> Candidate Gen Fast retrieval narrows the set to hundreds or thousands.
- Candidate Gen -> Ranker A stronger model scores candidates with more features.
- Ranker -> Reranker A final precise stage handles diversity, freshness, policy, or LLM reranking.
- Reranker -> Top Results Each stage trades recall, precision, latency, and business constraints.
In an interview
Multistage ranking is a funnel: you can't run an expensive model on millions of items, so cheap candidate generation retrieves thousands, a mid-weight ranker scores them with more features, and a final reranker handles precision, diversity, and business rules. Early stages optimize recall and speed; later stages optimize precision.
Production defaults
- Funnel
- millions → ~1000 (cheap retrieval) → ~100 (ranker) → ~10 (heavy reranker)
- Budget
- the whole funnel under your latency target; the last stage gets the most compute per item
- Measure per stage
- recall@k at each boundary. A late-stage fix can't recover an early-stage miss
What breaks
- Great ranker, poor results — The right item never made the candidate set. Measure recall at the retrieval stage first.
- Latency spikes at p95 — Late-stage depth. Cap candidates entering the expensive stage rather than optimizing the model.