发表机构
UC Berkeley; Microsoft Research(加州大学伯克利分校; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明测试时通信能显著提升多智能体在难题上的表现,团队成功率相当于四倍独立智能体,且优势随规模增长,并迁移至研究任务,超越已知最佳人类方案。
AI 中文摘要
科学进步并非孤立发生,而是通过协作实现,然而现有的智能体系统对此鲜有体现。通信智能体是否有帮助仍是一个悬而未决的问题,且先前结果不一。我们表明,在具有挑战性的任务上,测试时通信可以显著优于独立的并行尝试,其中共享突破性进展可以推动整个团队前进。我们首先研究了多智能体测试时通信扩展的影响,其中智能体没有预定义角色,通过共享目录进行通信,并在ARC-AGI-3(一个需要新颖问题解决的基准)上进行评估。我们发现,由k个通信智能体组成的团队(team@k)的成功率与4k个独立智能体相当,且这一优势随k增大而增长,表明收益随规模复合增长。这种效果不仅仅是效率问题:一个没有任何单个智能体能解决的任务,一个智能体团队可以可靠地解决。此外,在计算资源充足的情况下,这些收益可迁移到研究导向的任务上。在多联骨牌填充中,通信智能体优于best@k,并超过了先前已知的最佳分数。在MNIST分类器压缩中,通信超越了已知最佳的人类解决方案。一个由四个智能体组成的团队生成了一个1,957字节的分类器提交,达到了99.4%的测试准确率,比已知最佳的人类解决方案和最佳单智能体结果都更小。这些收益并非无条件。当计算资源有限或缺乏明确的进展衡量标准时,独立智能体可能优于通信。然而,在充足的计算资源和清晰的反馈下,多智能体通信始终能产生更强的结果。
英文摘要
Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of $k$ communicating agents, team@$k$, matches the success rate of $4k$ independent agents, and this advantage grows with $k$, suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@$k$ and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.
Comments34 pages, 12 figures