arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

协作损耗:LLM多智能体系统协调时的性能损失有多大

The Collaboration Tax: How Much LLM Multi-Agent Systems Pay to Coordinate

Weixiang Sun, Zehong Wang, Hong Huang, Colby Nelson, Yijun Ma, Yanfang Ye

arXiv 2608.22152首次发表:更新:

AI 中文总结

该研究定义了LLM多智能体系统的协作损耗,揭示其随模型能力提升单调下降,可通过对话特征预测,提示干预可缩小差距,为优化多智能体协作提供了量化依据。

AI 中文摘要

由大语言模型(LLM)构建的多智能体系统已被广泛部署,但当两个LLM必须协调而非单独行动时,性能会损失多少仍不清楚。我们将协作损耗定义为具有私人信息的两人合作博弈的团队去中心化损失,提出两个命题来表征其符号及其与最大超加性违反的等价性。我们将该定义应用于32个可单独处理的任务,这些任务按接地摩擦源分组,并在来自7家提供商的11个模型上对其进行测量。该损耗沿两个无例外的轴呈现:所有模型间的类别排序,以及随能力提升呈单调下降。直接机制并非推理缺陷,而是四阶段对话级联:智能体提出无根据的主张、未能查询伙伴、跳过整合双方观点、未重新推导便接受答案。该损耗可通过对话特征机械预测,且部分可处理:针对所有四个阶段的提示干预可缩小相当一部分差距,且不同类别的主要瓶颈不同。在异质配对中,该损耗向更强的伙伴倾斜而非加性中点,实证实现了我们框架预测的最大超加性违反。综上,这些结果将LLM系统中的协作重新定义为可测量、可预测且部分可处理的成本。

英文摘要

Multi-agent systems built from large language models are deployed widely, yet how much performance is lost when two LLMs must coordinate rather than act alone remains unclear. We formulate the collaboration tax as the team-decentralisation loss of a two-player cooperative game with private information, with two propositions characterising its sign and its equivalence to a max-superadditivity violation. We operationalise this definition on 32 solo-tractable tasks grouped by source of grounding friction and measure it on 11 models from 7 providers. The tax is structured along two no-exception axes: a category ordering across every model and a monotonic decrease with capability. The proximate mechanism is not a reasoning deficit but a four-stage conversational cascade in which agents make ungrounded claims, fail to query the partner, skip integrating both views, and accept the answer without re-derivation. The tax is mechanically predictable from conversation features and partly tractable: a prompt intervention targeting all four stages closes a substantial fraction of the gap, with the dominant bottleneck differing across categories. In heterogeneous pairs the tax is pulled toward the stronger partner rather than the additive midpoint, empirically realising the max-superadditivity violation predicted by our framework. Together these results recast collaboration in LLM systems as a measurable, predictable, and partly tractable cost.

CommentsEMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑