arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估上下文协议(ECP):一种用于AI智能体评估的便携式契约

The Evaluation Context Protocol (ECP): A Portable Contract for AI Agent Evaluation

Aniket Wattamwar, Manav Anandani, Mrunal Kakirwar

arXiv 2608.19263首次发表:更新:

AI 中文总结

针对当前AI智能体评估范式的局限性,提出厂商中立的ECP协议,实现跨框架的统一评估,已开发适配主流智能体框架的开源参考实现,相关验证为未来工作。

AI 中文摘要

人工智能的发展已要求从评估孤立的大型语言模型(LLMs)向评估自主智能体架构发生根本性转变。本文探讨了评估AI智能体的关键方法论以及高级可观测性基础设施的核心作用。我们分析了智能体的架构组件,并指出了当前评估范式的严重局限性,包括基准利用、“自信地错误”现象,以及理论能力与操作可靠性之间的差异。为解决当前评估基础设施的碎片化问题,本文提出了评估上下文协议(ECP),这是一个早期阶段、厂商中立的框架,旨在作为智能体系统的便携式评估契约层。目前,ECP定义了一个小型JSON-RPC接口,智能体通过该接口暴露其用户可见的输出、所做的工具调用以及评估器安全的审计上下文,程序检查可在各框架和持续集成系统上针对该接口统一运行。我们描述了一个开源参考实现,其中包含LangChain、LlamaIndex、CrewAI和PydanticAI的适配器,并将该设计与近期文献中记录的失败模式相结合。ECP被视为一项进行中的工作,而非已完成的标准:随着该协议在更多系统上应用,评估面、方法集和 grader 类别预计都会发生变化,而证明采用该协议所需的实证验证被列为未来工作。

英文摘要

The evolution of artificial intelligence has necessitated a fundamental shift from evaluating isolated Large Language Models (LLMs) to assessing autonomous agentic architectures. This paper explores the critical methodologies for evaluating AI agents and the essential role of advanced observability infrastructure. We analyze the architectural components of agents and identify the severe limitations of current evaluation paradigms, including benchmark exploitation, the "confidently wrong" phenomenon, and the discrepancy between theoretical capability and operational reliability. To begin addressing the fragmentation in current evaluation infrastructure, this paper proposes the Evaluation Context Protocol (ECP), an early-stage, vendor-neutral framework intended to act as a portable evaluation contract layer for agentic systems. In its current form ECP defines a small JSON-RPC interface over which an agent exposes its user-visible output, the tool calls it made, and evaluator-safe audit context, and against which programmatic checks can be run uniformly across frameworks and continuous integration systems. We describe an open-source reference implementation that includes adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI, and we situate the design against failure modes documented in the recent literature. ECP is presented as work in progress rather than a finished standard: the evaluation surface, method set, and grader families are all expected to change as the protocol is exercised against more systems, and the empirical validation required to justify adoption is outlined as future work.

Comments14 pages, 4 tables, 4 figures, Code available at https://github.com/evaluation-context-protocol

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑