CoCoBench:用于具身多智能体任务规划的协作协调基准
CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning
浏览论文内容
中文总结 AI 辅助
研究人员提出CoCoBench这一家庭任务多智能体具身协调基准,评估11种MLLM后发现协调能力具构造特异性,为相关模型设计提供新方向。
中文摘要 AI 辅助
近年来,由多模态大语言模型(MLLM)驱动的智能体系统发展迅速,但现有的具身智能体基准仍缺乏针对多智能体协调的细粒度诊断。大多数基准要么聚焦于单智能体任务完成,要么通过整体任务成功率来总结多智能体行为,这可能会掩盖诸如重复工作、顺序约束违反、资源竞争以及交接不同步等协调失败问题。本文中,我们提出了CoCoBench,这是一个用于评估可执行家庭任务中多智能体具身协调的构造级基准。CoCoBench包含897个经预言机验证的实例,围绕四个反复出现的协调构造组织:任务分配、顺序执行、互斥和交接协调。除任务成功率外,CoCoBench还提供构造级分数,用于衡量智能体是否有效协调。我们在不同协调模式、观测输入和智能体数量下评估了11种领先的MLLM。结果表明,协调能力具有高度构造特异性:整体性能强劲并不意味着在不同协调类型上具备均衡的能力。这些发现为设计针对性的模型架构和提升多智能体协调能力指明了新方向。
英文摘要
Agent systems powered by multimodal large language models (MLLMs) have advanced rapidly in recent years, yet existing embodied-agent benchmarks still lack fine-grained diagnostics for multi-agent coordination. Most benchmarks either focus on single-agent task completion or summarize multi-agent behavior with overall task success rates, which can obscure coordination failures such as duplicated work, violations of ordering constraints, resource contention, and desynchronized handoffs. In this paper, we introduce CoCoBench, a construct-level benchmark for evaluating multi-agent embodied coordination in executable household tasks. CoCoBench contains 897 oracle-validated instances organized around four recurring coordination constructs: task allocation, sequential ordering, mutual exclusion, and handoff coordination. In addition to task success rate, CoCoBench provides construct-level scores that measure whether agents coordinate effectively. We evaluate 11 leading MLLMs across different coordination modes, observation inputs, and numbers of agents. The results show that coordination ability is highly construct-specific: strong overall performance does not imply balanced competence across different coordination types. These findings point to new directions for designing targeted model architectures and improving multi-agent coordination ability.
发表机构
- Nanjing University(南京大学)
- AgiBot
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。