Let the model decide whether to retrieve, then grade its own evidence and its own answer — retrieval on demand instead of retrieval on reflex.
A careful researcher who asks three questions before answering: do I even need to look this up, is what I found actually about the question, and does my sentence really follow from it?
The two failure modes it removes — retrieving when you shouldn't, and generating past your evidence — are the two that destroy trust fastest.
Standard RAG retrieves for every query, whether or not retrieval helps, and then trusts whatever comes back. Self-RAG makes both decisions explicit and learned. The model emits reflection tokens: Retrieve? decides whether external evidence is needed at all; IsRel grades each retrieved passage for relevance; IsSup checks whether the draft sentence is actually supported by the passage it cites; IsUse rates overall usefulness. Generation branches over passages and the critique scores select the continuation, so unsupported claims lose to supported ones during decoding rather than being caught afterwards. The practical descendant is the graded RAG loop — relevance grader, groundedness check, answer check — which most teams implement with a small model rather than special tokens.
Self-RAG adds three decisions to the RAG loop that normal pipelines leave implicit: whether to retrieve at all, whether each retrieved passage is relevant, and whether the generated sentence is actually supported by it. In the original work these are learned reflection tokens that steer decoding; in practice most teams implement the same shape with a small grader model. The payoff is fewer unsupported claims and less pointless retrieval.
THE SELF‑RAG FRAMEWORK – RETRIEVAL TOKENS & REFLECTION — Sujatha Mudadla EduTech, 3:37