arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自组织智能体团队学会共同推理

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

arXiv 2609.22682首次发表:更新:

发表机构

Stanford University; Together AI; Goizueta Business School, Emory University(斯坦福大学; Together AI; 埃默里大学戈伊苏埃塔商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出自组织智能体团队(SAT),通过从少量协作中学习可复用策略来组织角色与信息流,在数学和物理基准上显著超越最强成员,并发现可证明性驱动改进。

AI 中文摘要

集体智能不仅取决于团队成员的知识,还取决于他们如何组织工作。当解决方案的结构未知时,无法预先指定有用的角色和分工;团队必须从经验中学习如何在推理展开时组织推理。人类团队通常会以这种方式适应,而现有的人工智能智能体团队则依赖固定协议、显式任务分解或路由。我们引入了自组织智能体团队(SAT),即固定的人工智能智能体团队,它们从先前的协作中学习可复用的策略,以组织角色、对话阶段、参与度和信息流。这些策略实现了我们所谓的协作计算:智能体交换、挑战、修复并综合部分推理,形成没有任何成员独立产生的解决方案。在两个独立的设置中,我们学习了团队合作策略,这些策略无需修改即可转移到未见过的基准测试,仅使用15个数学问题和25个研究生水平知识问题。在五个数学和物理基准测试中,自组织团队的平均准确率为66.7%,而其最强成员的准确率为48.8%,该成员的计算匹配推理准确率为58.7%,完美路由器对成员独立答案的准确率为59.0%;在AIME 2026上,它们比该路由器高出13.4个百分点。由于收益因基准而异,我们询问自组织协作何时有帮助。在八个基准测试中,可证明性(组织心理学中关于团队能否区分正确与错误推理的构念)与相对最强成员的改进强烈相关(Spearman ρ=0.90,p=0.005):当正确推理一旦出现就能被识别时,团队受益最大。更广泛地说,这些结果表明,组织本身可以成为一种智能体能力:智能体团队可以学习如何共同推理,并产生其成员无法独立达成的解决方案。

英文摘要

Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasoning as it unfolds. Human teams routinely adapt this way, while existing AI agent teams rely on fixed protocols, explicit task decomposition, or routing. We introduce Self-Organizing Agent Teams (SAT), fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow. These strategies enable what we call collaborative computation: agents exchange, challenge, repair, and synthesize partial reasoning into solutions no member produced independently. In two independent settings, we learn teamwork strategies that transfer unchanged to unseen benchmarks, using only 15 mathematics and 25 graduate-level knowledge problems. Across five mathematics and physics benchmarks, self-organizing teams average 66.7% accuracy, versus 48.8% for their strongest member, 58.7% for compute-matched inference by that agent, and 59.0% for a perfect router over members' independent answers; on AIME 2026, they exceed this router by 13.4 points. Because gains vary across benchmarks, we ask when self-organizing collaboration helps. Across eight benchmarks, demonstrability (the organizational-psychology construct of whether a team can distinguish correct from incorrect reasoning) strongly tracks improvement over the strongest member (Spearman $ρ=0.90$, $p=0.005$): teams benefit most when correct reasoning can be recognized once it appears. More broadly, these results suggest that organization itself can become an agent capability: agent teams can learn how to reason together and produce solutions their members could not reach independently.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑