All concepts

Multistage Ranking

Rank items in stages: broad retrieval, lightweight scoring, heavy reranking, and final business rules.

Applied ML · Advanced · ~8 min

In plain English

Millions of candidates, one screen of results. Narrow it in stages: something crude and instant first, something expensive and accurate at the end.

Why it's worth your time

It's the architecture of every search, feed and recommendation system — and the same shape as retrieve-then-rerank in RAG.

If you remember three things

  • Retrieval → ranking → re-ranking, each smaller and more expensive
  • Recall is the earlier stage's job; precision is the later stage's
  • A candidate lost early can never be recovered

Overview

A funnel that ranks huge candidate sets in stages: cheap retrieval narrows millions to thousands, a stronger ranker scores those with richer features, and a precise reranker applies diversity, freshness, and policy. Each stage trades recall for precision so heavy models only touch a small set.

How it works

  1. Start: Huge Corpus Millions of candidates cannot all go through an expensive model.
  2. Huge Corpus -> Candidate Gen Fast retrieval narrows the set to hundreds or thousands.
  3. Candidate Gen -> Ranker A stronger model scores candidates with more features.
  4. Ranker -> Reranker A final precise stage handles diversity, freshness, policy, or LLM reranking.
  5. Reranker -> Top Results Each stage trades recall, precision, latency, and business constraints.

In an interview

Multistage ranking is a funnel: you can't run an expensive model on millions of items, so cheap candidate generation retrieves thousands, a mid-weight ranker scores them with more features, and a final reranker handles precision, diversity, and business rules. Early stages optimize recall and speed; later stages optimize precision.

Production defaults

Funnel
millions → ~1000 (cheap retrieval) → ~100 (ranker) → ~10 (heavy reranker)
Budget
the whole funnel under your latency target; the last stage gets the most compute per item
Measure per stage
recall@k at each boundary. A late-stage fix can't recover an early-stage miss

What breaks

  • Great ranker, poor results — The right item never made the candidate set. Measure recall at the retrieval stage first.
  • Latency spikes at p95 — Late-stage depth. Cap candidates entering the expensive stage rather than optimizing the model.

Watch it explained

Data Science and Recommender Systems with Andrew | 365 Data Use Cases — 365 Data Science, 5:57

Related