All concepts

Agentic RAG Controller

Retrieval as a controlled loop, not a function call: route, decompose, retrieve per sub-question, grade, escalate, and stop on a budget.

Advanced RAG · Advanced · ~8 min

In plain English

A good analyst doesn't run one search. They work out what the question is really asking, split it up, look each part up in the right place, notice when an answer is missing, and stop when they have enough.

Why it's worth your time

Every question with two parts — a comparison, a filter and a lookup, arithmetic over retrieved numbers — is unanswerable in one retrieve-then-generate hop.

If you remember three things

  • Route first: some questions need no corpus at all
  • Decompose only what needs decomposing
  • Budget the loop, keep provenance, define the failure answer

Overview

Classic RAG is one retrieve-then-generate hop, which is the wrong shape for any question with more than one part. An agentic RAG controller treats retrieval as a policy. A router first decides whether the question needs the corpus, a SQL table, a web search, or nothing. A planner decomposes multi-part questions into sub-questions that can be answered independently. Each sub-question runs its own hybrid retrieval and reranking. A grader checks coverage and either accepts, rewrites, or escalates to a different source. Evidence is accumulated with provenance, and the synthesiser answers only from what survived. What makes it production-grade rather than a demo is the control layer: a step budget, a token budget, deduplicated evidence, and a defined answer for 'we couldn't establish this'.

In an interview

An agentic RAG controller replaces the single retrieve-then-generate hop with a loop that has a policy. It routes the question to the right source, decomposes multi-part questions, retrieves and grades per sub-question, rewrites or escalates when evidence is thin, and synthesises only from what passed — all under an explicit step and token budget. The budget and the refusal path are what make it shippable rather than a demo.

Production defaults

Budget
6 retrieval steps and a hard token cap per request; emit a partial answer with evidence when the cap hits
Routing
a small classifier or a tool-choice call — skipping retrieval on follow-ups is the cheapest latency win in the whole pipeline
Tracing
one span per stage (route, sub-question, retrieve, grade, synthesise) with ids. Without it, failures are guesswork

What breaks

  • One query burned the token budget — No step cap. Bound the loop and return what you have with an honest note about what's missing.
  • Context full of near-duplicates — No evidence deduplication. Three copies of one paragraph crowd out the one contradicting fact.
  • Can't tell why an answer was wrong — Provenance was dropped at synthesis. Carry source ids through to the final claim.

Watch it explained

What is Agentic RAG? — IBM Technology, 5:41

Related