A planner, researcher, critic, and writer can look like a complete architecture before anyone has established that four agents help. Each handoff adds messages, state, latency, and a chance to lose a constraint. More roles create more activity; they do not automatically create a better result.

Towards a Science of Scaling Agent Systems, revised in April 2026, compares agent architectures under controlled conditions and finds that task structure matters. Some decomposable tasks benefit from coordination; sequential and tool-heavy work can incur harmful overhead. These are benchmark-conditioned findings, not a rule that one topology always wins.

Decomposition is a hypothesis

Consider a synthetic supplier brief. Searching independent approved source collections may parallelize well. Choosing the final supplier may not: it depends on the same budget, qualification rules, and conflicting evidence. Splitting that decision among agents can create incompatible assumptions rather than useful diversity.

Draw the dependency graph before choosing roles. Identify independent work, shared resources, ordering constraints, and the person or component that resolves disagreement. If the graph is mostly a chain, parallel agents may spend more time synchronizing than solving.

Give the single-agent baseline equal tools

A weak baseline can manufacture an architecture win. Give the single-agent candidate access to the same tools, evidence, policy, and appropriate compute budget. Compare a deterministic workflow, one capable agent, and the proposed multi-agent arrangement on identical initial states.

Compute parity requires care. Equal numbers of calls are not equal cost if model sizes and context lengths differ. Report both task results and the actual resource budget. An architecture improvement bought by additional compute can still be useful, but it should be named.

Inspect the seams

For each handoff, record the task, required evidence, constraints, output format, and success condition. A researcher returning “approved supplier” without the applicable qualification version has produced a misleading local success. The final writer should not infer the missing approval from a fluent summary.

Use structured messages where they clarify the contract, and preserve source identifiers through summaries. Shared state needs ownership: decide who can revise a conclusion, invalidate evidence, or stop the workflow. Independent agents reading each other's unverified conclusions can amplify an early mistake.

Coordination quality has its own measures

MultiAgentBench evaluates collaboration and competition with milestone-based indicators and different coordination protocols. It illustrates why final task success alone does not describe interaction quality. A useful local suite can record redundant tool calls, conflicting intermediate claims, lost constraints, and time spent waiting for other components.

Those measures are diagnostics. Low message count is not the final goal, and identical conclusions can reflect shared bias rather than independent confirmation. Use end-state validators and expert review for conclusions that matter.

Run ablations before adding another agent

Remove the critic and see which failures return. Replace an agent's classification step with a rule. Collapse two sequential roles into one. Give the writer direct access to original passages rather than only a compressed handoff. These experiments reveal which components contribute and which merely rename existing work.

Repeat stochastic trials and keep failure severity separate. A multi-agent system that raises mean success but adds an unauthorized action has not automatically earned release. Examine degraded operation when one specialist is unavailable.

Choose the smallest useful coordination contract

Add a specialist when it contributes distinct tools, independent evidence, a validated check, or work that can truly proceed separately. Keep the coordinator's authority and stop conditions explicit. If the benefit disappears against a fair baseline, retain the simpler workflow.

Architecture should follow the evidence about the task. The interesting research question is which coordination mechanism produces a repeatable improvement under the operating constraints, rather than how many agents fit in the diagram.

Sources and further reading