Single agent by default; add an orchestrator, a handoff, or parallelism only when a specific constraint forces it.
Choosing the shape of the work: a straight pipeline, a branch, a loop, or several specialists. Pick the simplest one that can express the task.
Most agent complexity is self-inflicted — a fixed workflow would have done the job, faster and testably.
The default architecture is one agent with one context, and it wins more often than the diagrams suggest: nothing is lost in a handoff, no worker duplicates work, and one trace explains the whole run. Orchestration is what you reach for when a specific constraint appears. Independent, read-only subtasks that dominate wall-clock justify an orchestrator fanning out to workers — at the cost of roughly N× the tokens and a merge step. Distinct specialisations or permission boundaries justify routing, where control transfers with the state rather than splitting. And dependencies decide sequencing: if step N needs step N−1's output, no amount of concurrency makes it correct. The engineering judgment is knowing which constraint you actually have.
I start with one agent and one context, because nothing is lost in a handoff and one trace explains the run. I add orchestration only for a named constraint. Independent, read-only subtasks that dominate wall-clock justify an orchestrator fanning out to workers — but that costs roughly N× the tokens plus a merge step, so it's money for latency. Specialisation or a permission boundary justifies routing, where control transfers and the handoff must carry the goal, what's been tried, the evidence and the remaining budget. Dependencies decide sequencing: if step N needs N−1's output it stays a chain. Otherwise, one well-built agent with sharp tools and real evals beats the mesh.
Conceptual Guide: Multi Agent Architectures — LangChain, 8:58