arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30661cs.CL

SwarmBench:大语言模型能否充当智能体群的编排者?

SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?

发表机构中国科学院自动化研究所 · 中国科学院大学人工智能学院
查看机构详情
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

Jinshan Gao, Zhuoran Jin, Tianyi Men, Kang Liu, Jun Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大语言模型作为智能体群编排者的评估需求,提出多维度基准SwarmBench,发现模型编排能力存在显著差异,并通过SwarmExp方法有效提升了编排性能。

中文摘要 AI 辅助

基于大语言模型的多智能体系统正从固定交互拓扑向动态编排的Agent Swarm(智能体群)演进。然而,现有基准大多基于单智能体或通用智能体任务,难以系统评估关键编排能力。我们提出SwarmBench,一个从准确率、效率、成本和过程质量等多维度评估模型性能的基准。实验结果显示,当前模型的编排能力存在显著差异,这些差异不仅体现在最终的准确率、效率和成本上,也体现在编排过程的整体质量上。基于这些发现,我们进一步提出SwarmExp,一种基于经验提取和经验回放的简单但有效的方法,可持续提升大语言模型的编排性能。

英文摘要

Large language model-based multi-agent systems are evolving from fixed interaction topologies toward dynamically orchestrated Agent Swarms. However, existing benchmarks are still largely based on single-agent or general-purpose agent tasks, making it difficult to systematically evaluate key orchestration capabilities. We propose SwarmBench, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality. Experimental results show that current models exhibit substantial differences in orchestration capability. These differences are reflected not only in final accuracy, efficiency, and cost, but also in the overall quality of the orchestration process itself. Based on these findings, we further propose SwarmExp, a simple yet effective method based on experience extraction and experience replay, which consistently improves the orchestration performance of large language models.

补充信息

↑