发表机构
Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型演变成自主智能体后评估基础设施分散的问题,提出AgentCompass,通过围绕三个独立组件组织评估过程,具备容错异步运行时和轨迹分析工具,支持多基准测试,为智能体研究提供可扩展、可重复的基础设施。
AI 中文摘要
随着大语言模型(LLMs)演变成自主智能体,对统一评估基础设施的需求变得至关重要。然而,当前评估流程高度分散且紧密耦合,阻碍了可重复性并导致冗余工程。为解决此问题,我们引入了AgentCompass,这是一个用于评估基于LLM的智能体的开源、轻量级且可扩展的基础设施。AgentCompass围绕三个独立组件(即基准测试、测试工具和环境)组织评估过程,无需重新实现复杂执行逻辑即可实现灵活配置。此外,它具有容错异步运行时和全面的轨迹分析工具,可透明地诊断细微的失败模式,如奖励破解。AgentCompass原生支持五个能力维度上的20多个基准测试,为社区提供了一个用于推进智能体研究的可扩展且可重复的基础设施。
英文摘要
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.