arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentCompass:用于智能体能力的统一评估基础设施

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

Kai Chen, Zichen Ding, Jiaye Ge, Shufan Jiang, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tianhao Liang, Shudong Liu, Zerun Ma, Zixin Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu

arXiv 2607.13705首次发表:更新:

发表机构

Shanghai AI Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型演变成自主智能体后评估基础设施分散的问题,提出AgentCompass,通过围绕三个独立组件组织评估过程,具备容错异步运行时和轨迹分析工具,支持多基准测试,为智能体研究提供可扩展、可重复的基础设施。

AI 中文摘要

随着大语言模型(LLMs)演变成自主智能体,对统一评估基础设施的需求变得至关重要。然而,当前评估流程高度分散且紧密耦合,阻碍了可重复性并导致冗余工程。为解决此问题,我们引入了AgentCompass,这是一个用于评估基于LLM的智能体的开源、轻量级且可扩展的基础设施。AgentCompass围绕三个独立组件(即基准测试、测试工具和环境)组织评估过程,无需重新实现复杂执行逻辑即可实现灵活配置。此外,它具有容错异步运行时和全面的轨迹分析工具,可透明地诊断细微的失败模式,如奖励破解。AgentCompass原生支持五个能力维度上的20多个基准测试,为社区提供了一个用于推进智能体研究的可扩展且可重复的基础设施。

英文摘要

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑