arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ForestBench:用于评估多智能体协作的统一图框架

ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration

Guo Chen, Ziwen Li, Reed Li, Yu Lu, Haibo Shi, Bingbing Xu, Junjie Huang

arXiv 2608.08605首次发表:更新:

发表机构

Southwest University; Tencent; Institute of Computing Technology, Chinese Academy of Sciences(西南大学; 腾讯; 中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对多智能体系统异构轨迹评估的问题,提出ForestBench框架,将MAS轨迹映射为统一协作图,筛选844个查询并预计算参考图,可快速评估MAS协作。

AI 中文摘要

基于大语言模型(LLM)构建的多智能体系统(MAS)正迅速发展,但它们的异构执行轨迹无法为跨方法评估提供共同基础。仅基于结果的基准会忽略协作,而用LLM作为评判者的评估需要额外的、依赖模型的推理,且会随LLM和评分标准变化。我们提出一种可泛化的评估框架,将原生MAS轨迹映射到统一协作图的共享空间,使不同方法能在相同表示、参考集和指标面板下评估。候选图与特定查询的参考森林比较,每个森林是基准提供的经验证成功图的集合,记录代表性MAS方法完成任务的多样方式,而非规定唯一最优流程。将该框架实例化为ForestBench后,我们从7个公开数据集中筛选出844个需协作的查询,为每个查询预计算10个成功的目标条件参考图,并评估6个代表性MAS框架。通过控制主干、参考构建和扰动研究测试评估的稳定性与范围。一旦基准森林构建完成,ForestBench可在数毫秒内对轨迹评分,无需额外LLM推理,为比较不同MAS协作轨迹提供可复用的结构基础。

英文摘要

Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across methods. Outcome-only benchmarks discard collaborations, whereas LLM-as-Judge evaluation requires additional, model-dependent inference and can vary with the LLM and rubric. We introduce a generalizable evaluation framework that maps native MAS traces into a shared space of unified collaboration graphs, enabling different methods to be evaluated under the same representation, reference set, and metric panel. Candidate graphs are compared with a query-specific reference forest. Each forest is a benchmark-provided collection of verified-success graphs: it records diverse ways in which representative MAS methods can complete the task, rather than prescribing a unique optimal process. Instantiating the framework as ForestBench, we filter $844$ collaboration-necessary queries from seven public datasets, precompute ten successful target-conditioned reference graphs per query, and evaluate six representative MAS frameworks. Controlled backbone, reference-construction, and perturbation studies test the stability and scope of evaluation. Once the benchmark forests are built, ForestBench scores a trace in milliseconds without further LLM inference, providing a reusable structural basis for comparing diverse MAS collaboration traces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑