All concepts
Corrective RAG
Grade retrieved chunks; if they're weak, rewrite the query or fall back to web search before answering.
RAG & Retrieval · Advanced · ~8 min
In plain English
Before answering, check whether what you retrieved is actually any good. If it isn't, search again differently instead of answering from bad material.
Why it's worth your time
It's the fix for the most damaging RAG failure — confidently answering from irrelevant passages.
If you remember three things
- Grade retrieved documents for relevance before generating
- On failure: rewrite the query, widen the search, or say you don't know
- Adds a round trip; spend it only where it pays
Overview
Corrective RAG (CRAG) adds a self-check: a lightweight grader scores whether retrieved chunks are relevant. If confidence is low, the system rewrites the query, retrieves again, or falls back to another source — reducing answers built on irrelevant context.
How it works
- Retrieve for the query Embed the query and pull the top-k chunks — same as ordinary RAG so far.
- Grade the retrieval A lightweight grader scores whether the retrieved chunks are actually relevant to the query.
- Confident → generate If the grade is high, go straight to grounded generation from the good context.
- Weak → rewrite the query If the chunks look irrelevant, don't answer from them — rewrite or expand the query instead.
- Fall back to web search Re-retrieve, or fall back to web search, to find better evidence than the first attempt.
- Generate from good context Now generate using the corrected, relevant context.
- Grounded, cited answer Return a cited answer — the self-check means bad retrieval no longer silently produces wrong answers.
In an interview
Corrective RAG grades retrieval quality before generating. If the retrieved chunks look irrelevant, it corrects course — rewriting the query, re-retrieving, or falling back to web search — instead of answering from bad context. It's a control loop that targets RAG's dominant failure mode: bad retrieval.
Production defaults
- Grader
- a cheap model or a reranker score threshold, not another full LLM call if you can avoid it
- Fallback order
- query rewrite → widen k → web/alternate source → refuse
- Cap retries
- one extra attempt. Two is rarely better and always slower
What breaks
- Latency doubled for everyone — You're grading every request. Trigger correction only when the top reranker score is below a threshold.
- It rewrites good queries into bad ones — The grader is too strict. Calibrate its threshold against labelled relevant/irrelevant pairs.