Capable language models can outgrow the benefits of collaboration
Nature Machine Intelligence, Published online: 24 July 2026; doi:10.1038/s42256-026-01268-y A controlled study of large language model agents across 260 configurations shows when multi-agent collaboration helps or hurts performance, and introduces a predictive model that selects the best architecture in 87% of held-out within-domain configurations.
Recent advancements in capable language models have led to the widespread deployment of agent-based systems capable of reasoning, planning, and utilizing tools to complete tasks. However, it remains unclear when multi-agent coordination surpasses the performance of a single powerful agent. To investigate this, we conducted a controlled experiment analyzing 260 configurations across six benchmarks, five architectures, and three language model families, holding task prompts, tools, and compute budgets constant while varying coordination structures and model capabilities.
We discovered a predictive model using empirical coordination metrics, revealing that the performance of a single agent typically serves as the most reliable predictor of whether coordination enhances or diminishes performance. Our findings suggest a practical capability-saturation threshold beyond which additional agents are unlikely to improve performance, with 94% accuracy in predicting the impact of multi-agent coordination on tasks such as SWE-bench Verified and Terminal-Bench.
Furthermore, we uncovered a baseline-scaled error amplification effect, supporting a taxonomy of failure modes. This quantitative framework offers valuable insights for architecture selection and estimating when multi-agent coordination is likely to improve performance or introduce overhead.
Written by urgent.news from Nature Machine Intelligence's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.