All concepts

Mixture of Agents

Several models answer, a later layer reads all their answers and writes a better one — ensembling for language, in layers.

Advanced Agentic Systems · Advanced · ~6 min

In plain English

A panel where everyone writes an answer, then everyone reads the panel's answers and writes a better one. Two rounds of that and the last draft beats every first draft.

Why it's worth your time

Models write measurably better answers when shown other models' attempts — even attempts weaker than their own.

If you remember three things

  • Proposers answer; the next layer synthesises rather than re-answers
  • Diversity of models matters more than picking the best one
  • n× tokens and layer-barrier latency — this is not for chat

Overview

Mixture-of-Agents stacks language models the way a neural network stacks layers. In the first layer, several different models — deliberately heterogeneous — answer the same prompt independently. Their responses are concatenated and handed to the next layer, whose models are asked not to answer afresh but to synthesise: reconcile, correct, and combine. Two or three layers of this, then a final aggregator produces the answer. The finding that made it notable is collaborativeness: a model produces a better response when shown other models' attempts, even when those attempts are individually weaker than its own output. The costs are equally notable — n× tokens per layer and latency equal to the slowest model in each layer — so it belongs on high-value, quality-critical work, not on a chat turn.

In an interview

Mixture-of-Agents runs several different models on the same prompt in parallel, then feeds all their answers to a next layer whose job is to synthesise rather than re-answer, repeating for a couple of layers before a final aggregation. Models demonstrably write better answers when they can see other attempts, including weaker ones. The price is n× tokens and layer-wise latency, so it's for high-value work, not chat.

Production defaults

Shape
3-4 heterogeneous proposers, 2 layers, strongest model as the final aggregator
Timeouts
per-layer wall clock. A layer is a barrier, so one slow provider stalls the whole run
When
value per answer clearly above ~10× token cost. Benchmark against one strong model first

What breaks

  • No better than a single model — Your proposers are too similar — often several samples from one model. Correlated errors don't cancel.
  • Unusable latency — Layers are barriers. Either cut to one layer, or move it off the interactive path entirely.

Watch it explained

AI Agents vs Mixture of Experts: AI Workflows Explained — IBM Technology, 9:27

Related