All concepts

Corrective RAG

Grade retrieved chunks; if they're weak, rewrite the query or fall back to web search before answering.

RAG & Retrieval · Advanced · ~8 min

In plain English

Before answering, check whether what you retrieved is actually any good. If it isn't, search again differently instead of answering from bad material.

Why it's worth your time

It's the fix for the most damaging RAG failure — confidently answering from irrelevant passages.

If you remember three things

  • Grade retrieved documents for relevance before generating
  • On failure: rewrite the query, widen the search, or say you don't know
  • Adds a round trip; spend it only where it pays

Overview

Corrective RAG (CRAG) adds a self-check: a lightweight grader scores whether retrieved chunks are relevant. If confidence is low, the system rewrites the query, retrieves again, or falls back to another source — reducing answers built on irrelevant context.

How it works

  1. Retrieve for the query Embed the query and pull the top-k chunks — same as ordinary RAG so far.
  2. Grade the retrieval A lightweight grader scores whether the retrieved chunks are actually relevant to the query.
  3. Confident → generate If the grade is high, go straight to grounded generation from the good context.
  4. Weak → rewrite the query If the chunks look irrelevant, don't answer from them — rewrite or expand the query instead.
  5. Fall back to web search Re-retrieve, or fall back to web search, to find better evidence than the first attempt.
  6. Generate from good context Now generate using the corrected, relevant context.
  7. Grounded, cited answer Return a cited answer — the self-check means bad retrieval no longer silently produces wrong answers.

In an interview

Corrective RAG grades retrieval quality before generating. If the retrieved chunks look irrelevant, it corrects course — rewriting the query, re-retrieving, or falling back to web search — instead of answering from bad context. It's a control loop that targets RAG's dominant failure mode: bad retrieval.

Production defaults

Grader
a cheap model or a reranker score threshold, not another full LLM call if you can avoid it
Fallback order
query rewrite → widen k → web/alternate source → refuse
Cap retries
one extra attempt. Two is rarely better and always slower

What breaks

  • Latency doubled for everyone — You're grading every request. Trigger correction only when the top reranker score is below a threshold.
  • It rewrites good queries into bad ones — The grader is too strict. Calibrate its threshold against labelled relevant/irrelevant pairs.

Watch it explained

What Is Agentic RAG? When Retrieval Becomes a Loop — [AI Stack 18] — The AI Stack, 7:28

Related