arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21527cs.LG

OpenMAS-GCom:面向图增强多智能体系统的诊断基准

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

Kairui Yang, Xunkai Li, Kaixiang Zhang, Minghao An, Zekai Chen, Yuxuan Ba, Rong-Hua Li

AI总结:

针对图增强多智能体系统性能归因困难的问题,提出OpenMAS-GCom诊断基准,通过受控干预评估组件影响,在29个数据集上测试17种配置,发现不同组件干预导致不同性能变化。

AI中文摘要:

图增强多智能体系统(G-MAS)通过通信图和角色分配来协调大型语言模型智能体,这些机制决定了智能体之间如何交换信息以及如何划分职责。然而,跨系统的最终得分比较同时混合了模型、通信模式、角色和计算成本的差异,使得性能差异难以归因于特定的通信结构、角色分配和信息流。为解决这一评估归因问题,我们引入了OpenMAS-GCom,一个通过受控干预来诊断这些组件如何影响G-MAS性能的基准。我们通过协作单元、通信链路、共享中间信息和执行规则来表示系统。OpenMAS-GCom将原始系统与通过改变一个组件而修改的版本进行比较,同时保持任务、模型、提示和预算限制不变。我们重连通信边、移除专家或评论家智能体、用错误内容替换中间消息,并在执行期间禁用工作者。该基准在六个领域的29个数据集上评估了17种单智能体、普通多智能体和图增强配置。我们添加了400个G-MAS-Complex任务,要求智能体结合来自多个文档的信息、解决冲突的记录,并返回带有来源标识符的指定值。实验表明,移除专家后的平均损失大于移除评论家后的平均损失;在错误消息和工作者故障下,尽管原始得分相似,但性能下降不同;在G-MAS-Complex上,不同配置实现了最高的准确率和每token准确率。

英文摘要:

Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score comparisons across systems combine differences in models, communication patterns, roles, and computation costs, making performance differences difficult to attribute to specific communication structures, role assignments, and information flows. To address this evaluation attribution problem, we introduce OpenMAS-GCom, a benchmark for diagnosing how these components affect G-MAS performance through controlled interventions. We represent systems through collaboration units, communication links, shared intermediate information, and execution rules. OpenMAS-GCom compares original systems with versions modified by changing one component while keeping tasks, models, prompts, and budget limits fixed. We rewire communication edges, remove specialist or critic agents, replace intermediate messages with incorrect content, and disable workers during execution. The benchmark evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains. We add 400 G-MAS-Complex tasks requiring agents to combine information from multiple documents, resolve conflicting records, and return specified values with source identifiers. Experiments show larger mean losses after specialist removal than after critic removal, different performance degradation under incorrect messages and worker failures despite similar original scores, and different configurations achieving the highest accuracy and accuracy per token on G-MAS-Complex.

↑