All concepts

Deep Research Agents

Plan, fan out subagents with separate context windows, verify every claim against a source, then synthesise — the shape behind every 'deep research' product.

Advanced Agentic Systems · Advanced · ~7 min

In plain English

Sending four researchers to different libraries with clear briefs, then writing the report from their notes rather than from every book they touched.

Why it's worth your time

Breadth costs context, and one window can't hold five lines of inquiry. Separate windows per subagent is the mechanism, not the parallelism.

If you remember three things

  • Subagents return compressed findings with source ids, never raw pages
  • Specific briefs: objective, format, boundaries — vague ones overlap
  • Citation is its own pass; writing and citing together cites the nearest source

Overview

A research question is not a retrieval question: it needs breadth, comparison, and a defensible answer. The pattern that works is an orchestrator with subagents. The lead agent decomposes the question into independent lines of inquiry and spawns a subagent per line, each with its OWN context window — which is the real point, since parallel search would otherwise fill one window with hundreds of pages of raw source. Subagents search, read, and return compressed findings with source ids. The lead synthesises, and a separate citation pass attaches every claim to a source, dropping claims that cannot be supported. Anthropic reported that this multi-agent shape substantially outperformed a single agent on breadth-first research, and that token spend — roughly an order of magnitude over chat — is what buys the difference.

In an interview

A deep research agent decomposes a question into independent lines of inquiry and runs a subagent on each, every subagent with its own context window so raw sources never crowd the orchestrator's. Subagents return compressed findings with source ids; the lead synthesises; and a separate citation pass binds each claim to a source and drops what can't be supported. The context isolation and the separate citation pass are what make it work.

Production defaults

Fan-out
3-5 subagents, scaled to question complexity. A flat fan-out is either wasteful or shallow
Budget
expect roughly 15× a chat interaction in tokens. Cap per request and stream progress
Durability
checkpoint after each subagent — these runs are long enough that a mid-run failure must not restart from zero

What breaks

  • Subagents duplicated each other's work — Briefs were vague. State the objective, the output format, and what NOT to cover, per subagent.
  • Citations point at the wrong source — You cited during synthesis. Run a dedicated citation pass claim by claim and drop what can't be supported.
  • Lead agent's context blew up — Subagents returned raw pages. They must return summaries plus source ids — that isolation is the whole design.

Watch it explained

Deep Research Max: A step change for autonomous research agents — Google for Developers, 3:03

Related