arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EMAS:通过证据引导的修正稳定多智能体系统演化

EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision

Chao Fei, Qingyi Si, Kaihua Liang, Yanghua Xiao, Panos Kalnis, Hongcheng Guo

arXiv 2608.07196首次发表:更新:

发表机构

King Abdullah University of Science and Technology (KAUST); Fudan University(阿卜杜拉国王科技大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EMAS 是一种不更新 LLM 参数、利用样本经验修正 MAS 拓扑与提示词的方法,在四个基准和两个 LLM 上提升了准确率并降低了 token 成本,表现优于多数基线方法。

AI 中文摘要

许多自动化多智能体系统设计方法会在初始设计阶段优化提示词和拓扑结构,之后将生成的系统原封不动地部署到后续样本中。这些样本的经验很少被整合为可复用的系统更新,而以准确率为导向的设计可能会产生较高的 token 成本。我们引入 EMAS(Evolving Multi-Agent System,演化多智能体系统),该方法利用这些经验在不更新大语言模型(LLM)参数的情况下修正多智能体系统(MAS)的拓扑结构和提示词,以提高准确率或降低成本。EMAS 将轨迹转换为结构化诊断,指定修正操作和目标;仅当同一诊断在多个样本中重复出现时,才生成候选修正,且仅当对当前 MAS 的配对验证满足相应接受标准时,才应用该修正。在四个基准和两个 LLM 上,EMAS 在两种主干模型中均达到最高的任务加权总体准确率,且在八分之六的模型-基准设置中表现最佳或并列最佳。在两个演化周期内,EMAS 在 Kimi-K2-6 和 Qwen3.6-27B 上的任务加权准确率分别实现了 6.30% 和 20.10% 的相对提升;在 MBPP 数据集上,使用 Qwen3.6-27B 时,EMAS 将准确率从 55.09% 提升至 89.12%,同时每个任务的 token 使用量降低了 62.2%。这些结果表明,EMAS 能够将新样本的经验转化为可复用的 MAS 拓扑结构和提示词更新。

英文摘要

Many methods for automated multi-agent system design optimize prompts and topologies during an initial design stage and then deploy the resulting system unchanged on subsequent samples. Experience from these samples is rarely consolidated into reusable system updates, while accuracy-oriented designs may incur high token costs. We introduce EMAS (Evolving Multi-Agent System), which uses this experience to revise MAS topology and prompts without updating LLM parameters, either to improve accuracy or to reduce cost. EMAS converts traces into structured diagnoses that specify a revision operation and target. It generates a candidate revision only when the same diagnosis recurs across samples and applies it only if paired validation against the current MAS meets the corresponding acceptance criterion. Across four benchmarks and two LLMs, EMAS attains the highest task-weighted overall accuracy for both backbones and is best or tied in six of eight model--benchmark settings. Within two evolution epochs, EMAS achieves relative gains of 6.30% and 20.10% in task-weighted accuracy on Kimi-K2-6 and Qwen3.6-27B, respectively. On MBPP with Qwen3.6-27B, EMAS raises accuracy from 55.09% to 89.12% while reducing token use per task by 62.2%. These results show that EMAS can turn experience from new samples into reusable updates to MAS topology and prompts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑