Several models answer, a later layer reads all their answers and writes a better one — ensembling for language, in layers.
A panel where everyone writes an answer, then everyone reads the panel's answers and writes a better one. Two rounds of that and the last draft beats every first draft.
Models write measurably better answers when shown other models' attempts — even attempts weaker than their own.
Mixture-of-Agents stacks language models the way a neural network stacks layers. In the first layer, several different models — deliberately heterogeneous — answer the same prompt independently. Their responses are concatenated and handed to the next layer, whose models are asked not to answer afresh but to synthesise: reconcile, correct, and combine. Two or three layers of this, then a final aggregator produces the answer. The finding that made it notable is collaborativeness: a model produces a better response when shown other models' attempts, even when those attempts are individually weaker than its own output. The costs are equally notable — n× tokens per layer and latency equal to the slowest model in each layer — so it belongs on high-value, quality-critical work, not on a chat turn.
Mixture-of-Agents runs several different models on the same prompt in parallel, then feeds all their answers to a next layer whose job is to synthesise rather than re-answer, repeating for a couple of layers before a final aggregation. Models demonstrably write better answers when they can see other attempts, including weaker ones. The price is n× tokens and layer-wise latency, so it's for high-value work, not chat.
AI Agents vs Mixture of Experts: AI Workflows Explained — IBM Technology, 9:27