发表机构
Beijing University of Posts and Telecommunications; Shanghai Jiao Tong University; Tsinghua University(北京邮电大学; 上海交通大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有多智能体基准忽略部分可观测性的问题,提出MASBench基准,包含推理、调度、博弈三类任务,评估协议、记忆、路由三种协作机制,并提供性能、通信成本等确定性指标,指导MAS设计。
AI 中文摘要
大型语言模型(LLM)已逐步演变为自主智能体的核心。在此基础上,基于LLM的多智能体系统(MAS)将多个智能体协调成一个协同团队,以完成超出单个智能体能力的复杂任务。此类系统的有效性不仅取决于智能体本身,还取决于协作机制的设计和组织方式。注意,现实世界中的协作通常是部分可观测的,其中每个智能体由于物理或隐私相关的约束,只能访问环境的局部信息。然而,许多现有的多智能体基准假设全局可观测性,并且对系统评估协作机制的支持有限。为了弥合这一差距,我们引入了MASBench,一个在部分可观测约束下设计的多智能体协作基准。它被组织为三个递进的任务类别:推理、调度和博弈。通过这一结构,我们逐步评估三种代表性的协作机制:协议、记忆和路由。MASBench进一步提供了确定性评估指标,包括性能得分、通信成本和成本效益,以表征协作结果和通信开销。跨不同LLM主干和机制配置的实验为有效的MAS设计提供了经验指导。代码可在以下网址获取:此https URL
英文摘要
Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems depends not only on the agents themselves, but also on how collaboration mechanisms are designed and organized. Note that real-world collaboration is typically partially observable, where each agent can only access partial information about the environment due to physical or privacy-related constraints. However, many existing multi-agent benchmarks assume global observability, and leave limited support for systematically evaluating collaboration mechanisms. To bridge this gap, we introduce MASBench, a multi-agent collaboration benchmark designed under partially observable constraints. It is organized into three progressive task categories: Reasoning, Scheduling, and Game. Through this structure, we progressively evaluate three representative collaboration mechanisms: Protocol, Memory, and Routing. MASBench further provides deterministic evaluation metrics, including performance score, communication cost, and cost effectiveness, to characterize both collaboration outcomes and communication overhead. Experiments across diverse LLM backbones and mechanism configurations offer empirical guidance for effective MAS design. Code is available at: https://github.com/BUPT-GAMMA/MASBench