arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31590cs.MA

AgentWorld:多智能体大语言模型长时程协作基准

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Raphael Shu, Yusen Zhang, Young Min Cho, Jin Mo Yang, Yuan Yuan, Wenliang Zheng, Sharath Chandra Guntuku, Lyle Ungar, Zhou Yu, Rui Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

AgentWorld是一个包含100个长时程多智能体协作任务的基准,提出因果协作有效性(CCE)度量,实验显示最佳模型任务成功率仅52.0%,并揭示了通信中断等系统性失败模式。

中文摘要 AI 辅助

现有的多智能体基准主要测试竞争性场景、20步以内的短时程交互,或仅仅聚合个体表现,未能隔离并突出基于大语言模型的智能体的真实协作能力。我们提出AgentWorld,一个包含100个人工标注任务(附带100个增强变体)的基准,用于评估长时程、多智能体协作。任务跨越丰富的MMORPG沙盒环境中50多轮交互,需要3-20个具有不对称角色和能力的智能体通过通信、联合规划和资源共享进行协调,且处于黑盒设置中,每个智能体独立行动,无法访问其他智能体的内部状态。为了在常规二元任务成功之外量化协作有效性,我们提出因果协作有效性(CCE),一种基于图的度量,追踪智能体行为之间的因果依赖关系,并衡量团队努力中实际促成结果的比例。使用Gemini 3 Flash、Claude Haiku 4.5、GPT-5 Mini和DeepSeek R1-70B进行的实验表明,即使最佳模型也仅达到52.0%的任务成功率,系统性失败模式包括通信中断、角色混淆以及无法跨轮次维持共享计划。AgentWorld完全开源。

英文摘要

Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.

发表机构

  • OpenAgents
  • Columbia University(哥伦比亚大学)
  • University of Pennsylvania(宾夕法尼亚大学)
  • Seoul National University(首尔国立大学)
  • Penn State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑