arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34496cs.MA

MASTraceBench:通过基于LLM的多智能体系统中的提案轨迹诊断协作增益

MASTraceBench: Diagnosing Collaboration Gains through Proposal Trajectories in LLM-Based Multi-Agent Systems

Yapeng Li, Songze Li, Shuang Yu, Jing Yu, Zhixin Liu, Liqiang Wen, Tonghua Su

首次发表
浏览论文内容

中文总结 AI 辅助

针对MAS评估忽视协作增益过程的问题,提出MASTraceBench基准,通过提案轨迹诊断协作增益,并基于发现提出CLEARS方法,以声明级评估提升协作增益。

中文摘要 AI 辅助

基于LLM的多智能体系统(MAS)在复杂问题求解中展现出潜力。随着MAS方法的多样化,系统性评估变得越来越具有挑战性。然而,现有基准大多关注最终结果,导致协作增益如何产生、保持或丧失尚不清楚。为解决这一局限,我们引入MASTraceBench,一个通过MAS中的提案轨迹诊断协作增益的基准。在六个合作与竞争任务中,MASTraceBench跟踪并评分提案轨迹,并提供覆盖任务得分、协作增益、提案轨迹指标和令牌成本的多层指标套件。利用MASTraceBench,我们不仅按最终性能,还按智能体提案如何演变并聚合到最终答案来系统比较代表性MAS方法。分析揭示了一个反复出现的模式:最终MAS答案很少超越最强的初始提案;交互往往将较弱的初始提案提升至接近它,而强的初始提案很少被进一步改进,甚至可能退化。为降低这一风险,我们提出CLEARS,它用跨智能体的声明级评估取代整体提案交换,以指导可靠的综合。CLEARS更常保留或改进最强的初始提案,并在六个任务中的五个上取得最高协作增益。

英文摘要

LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this limitation, we introduce MASTraceBench, a benchmark for diagnosing collaboration gains through proposal trajectories in MAS. Across six cooperative and competitive tasks, MASTraceBench tracks and grades proposal trajectories and provides a multi-layer metric suite covering Task Score, Collaboration Gain, proposal-trajectory indicators, and Token Cost. Using MASTraceBench, we systematically compare representative MAS methods not only by final performance, but also by how agent proposals evolve and are aggregated into the final answer. This analysis reveals a recurring pattern: final MAS answers rarely surpass the strongest initial proposal; interaction often lifts initially weaker proposals toward it, while strong initial proposals are seldom further improved and may regress. To reduce this risk, we propose CLEARS, which replaces whole-proposal exchange with claim-level evaluation across agents to guide reliable synthesis. CLEARS more often preserves or improves upon the strongest initial proposal and achieves the highest Collaboration Gain on five of the six tasks.

发表机构

  • Harbin Institute of Technology(哈尔滨工业大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑