发表机构
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体适配框架的系统化评估难题,本文提出A²E引擎,借助ATP协议实现任务快速集成,通过多维指标评估适配框架能力,实验表明模型-适配框架组合性能差异显著,为模型与适配框架协同演进提供指导。
AI 中文摘要
随着大语言模型(LLMs)的快速发展,适配框架(harness)已成为在广泛领域部署智能体的关键基础设施。快速演变的适配框架生态系统也使得严格的能力评估愈发重要。然而,高效构建端到端、系统化且全面的评估流水线仍是一项重大挑战。为应对这一挑战,我们提出了A²E(Agent Auditing Engine,智能体审计引擎),这是一款专为智能体适配框架设计的端到端评估引擎。A²E利用我们新提出的智能体任务协议(Agent Task Protocol,ATP),实现了评估任务与不同适配框架的快速集成。通过自动插桩的监控器,它在实验过程中捕获并生成标准化的执行轨迹。在评估阶段,A²E使用一套多维指标系统地评估适配框架的能力。相较于仅评估正确性,这些指标能更细致地刻画适配框架在执行效率、工具使用、任务规划和错误恢复方面的差异。利用A²E开展的实验进一步表明,模型-适配框架组合在不同类型任务中表现出显著的性能差异,且没有任何一种组合能在所有任务中始终优于其他组合。这些发现不仅证明了系统化评估的必要性,还为模型与适配框架的协同演进提供了有用指导。我们的代码可在该https链接获取。
英文摘要
With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving harness ecosystem has also made rigorous capability evaluation increasingly important. However, efficiently building an end-to-end, systematic, and comprehensive evaluation pipeline remains a significant challenge. To address this challenge, we introduce $A^2E$ (Agent Auditing Engine), an end-to-end evaluation engine designed for agent harnesses. $A^2E$ leverages our newly proposed Agent Task Protocol (ATP) to enable the rapid integration of evaluation tasks with different harnesses. Through an automatically instrumented Monitor, it captures and generates standardized execution traces during experiments. In the Evaluation stage, $A^2E$ systematically assesses harness capabilities using a suite of multidimensional metrics. Compared with correctness alone, these metrics provide a more fine-grained characterization of differences among harnesses in execution efficiency, tool use, task planning, and error recovery. Experiments conducted with $A^2E$ further reveal that model-harness combinations exhibit substantial performance variation across different types of tasks, and that no single combination consistently outperforms all others across every task. These findings not only demonstrate the necessity of systematic evaluation but also provide useful guidance for the co-evolution of models and harnesses. Our code is available at https://github.com/datamllab/A2E.