All concepts

Self-RAG & Reflective Retrieval

Let the model decide whether to retrieve, then grade its own evidence and its own answer — retrieval on demand instead of retrieval on reflex.

Advanced RAG · Advanced · ~6 min

In plain English

A careful researcher who asks three questions before answering: do I even need to look this up, is what I found actually about the question, and does my sentence really follow from it?

Why it's worth your time

The two failure modes it removes — retrieving when you shouldn't, and generating past your evidence — are the two that destroy trust fastest.

If you remember three things

  • Retrieve? is a decision, not a reflex
  • Groundedness is the check that catches invention
  • A refusal with what it did find beats a fluent guess

Overview

Standard RAG retrieves for every query, whether or not retrieval helps, and then trusts whatever comes back. Self-RAG makes both decisions explicit and learned. The model emits reflection tokens: Retrieve? decides whether external evidence is needed at all; IsRel grades each retrieved passage for relevance; IsSup checks whether the draft sentence is actually supported by the passage it cites; IsUse rates overall usefulness. Generation branches over passages and the critique scores select the continuation, so unsupported claims lose to supported ones during decoding rather than being caught afterwards. The practical descendant is the graded RAG loop — relevance grader, groundedness check, answer check — which most teams implement with a small model rather than special tokens.

In an interview

Self-RAG adds three decisions to the RAG loop that normal pipelines leave implicit: whether to retrieve at all, whether each retrieved passage is relevant, and whether the generated sentence is actually supported by it. In the original work these are learned reflection tokens that steer decoding; in practice most teams implement the same shape with a small grader model. The payoff is fewer unsupported claims and less pointless retrieval.

Production defaults

Loop cap
2 rounds. A third almost never converges and always costs
Grader
a small model, all candidates in ONE batched call — per-document grading dominates latency
Failure mode
explicit refusal citing what was found. Make it a designed outcome, not an exception path

What breaks

  • Confident unsupported claims — You grade relevance but not groundedness. They are different checks and only the second catches invention.
  • Latency doubled — Serial per-document grader calls. Batch them, and skip retrieval entirely when the router says the turn doesn't need it.

Watch it explained

THE SELF‑RAG FRAMEWORK – RETRIEVAL TOKENS & REFLECTION — Sujatha Mudadla EduTech, 3:37

Related