Retrieval as a controlled loop, not a function call: route, decompose, retrieve per sub-question, grade, escalate, and stop on a budget.
A good analyst doesn't run one search. They work out what the question is really asking, split it up, look each part up in the right place, notice when an answer is missing, and stop when they have enough.
Every question with two parts — a comparison, a filter and a lookup, arithmetic over retrieved numbers — is unanswerable in one retrieve-then-generate hop.
Classic RAG is one retrieve-then-generate hop, which is the wrong shape for any question with more than one part. An agentic RAG controller treats retrieval as a policy. A router first decides whether the question needs the corpus, a SQL table, a web search, or nothing. A planner decomposes multi-part questions into sub-questions that can be answered independently. Each sub-question runs its own hybrid retrieval and reranking. A grader checks coverage and either accepts, rewrites, or escalates to a different source. Evidence is accumulated with provenance, and the synthesiser answers only from what survived. What makes it production-grade rather than a demo is the control layer: a step budget, a token budget, deduplicated evidence, and a defined answer for 'we couldn't establish this'.
An agentic RAG controller replaces the single retrieve-then-generate hop with a loop that has a policy. It routes the question to the right source, decomposes multi-part questions, retrieves and grades per sub-question, rewrites or escalates when evidence is thin, and synthesises only from what passed — all under an explicit step and token budget. The budget and the refusal path are what make it shippable rather than a demo.
What is Agentic RAG? — IBM Technology, 5:41