arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更糟的组合:多用户多智能体团队中的性能崩溃

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen

arXiv 2610.00583首次发表:更新:

发表机构

Stanford University; Anthropic(斯坦福大学; Anthropic)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨多用户多智能体团队协调失败问题,发现团队性能普遍不如单一协调者,并识别停滞、覆盖和捏造等行为,提出环境特定的缓解措施,发布MAMUBench基准。

AI 中文摘要

人们越来越多地将任务委托给AI智能体,而这些智能体越来越多地遇到其他人的智能体,它们共享代码库、日历或预算等资源。当每个智能体为具有不同目标的不同用户行动时,协调往往失败,最终群体的结果比单个智能体为所有人行动时更糟。我们在五个前沿模型和四个环境中的77个场景中研究了这种多用户、多智能体设置:一个API密钥环境(智能体共享计算预算)、一个诊所(共享日历)、一个个人助理环境(共享团体订单或预订)以及一个合并队列(共享发布截止时间)。在每个场景中,我们将一个为所有用户服务的单一智能体(协调者)与每个智能体为一个用户服务的团队进行比较,团队中智能体之间有无通信渠道两种情况。在每种环境中,团队提供的群体结果都比协调者差:没有渠道时,团队在两个环境中完全崩溃;即使有渠道,协调开销也会造成显著差距。例如,在个人助理环境中,协调者满足目标用户请求的频率约为团队的两倍。我们识别出与这种糟糕的群体级性能相关的不同行为,包括团队规模增大时的停滞、相互覆盖对方的行动以及捏造声明。我们找到了有效但特定于环境的缓解措施,例如团队负责人、明确的程序性指令以及平台检查(使智能体在提交前读取同伴的消息)。我们将发布API密钥、诊所和个人助理环境作为MAMUBench,包含74个场景,用于评估多用户、多智能体协调。

英文摘要

People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other's actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.

Comments63 pages, 22 Figures, 10 Tables, Code: https://github.com/safety-research/MAMUBench (will be released after review)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑