arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EIO-Agents:AI智能体评估缺失的语义层

EIO-Agents: The Missing Semantic Layer for AI Agent Evaluation

Fouad Bousetouane

arXiv 2610.07675首次发表:更新:

发表机构

The University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI智能体评估缺乏语义标准的问题,提出EIO-Agents规范,通过EIO本体和PER记录提供从证据到决策的可验证链条,使评估成为可独立核查的工件。

AI 中文摘要

AI智能体正进入日益重要的生产环境,但对其评估结果的实际含义缺乏统一的语义标准。分数、轨迹、评判输出和多陪审员发现越来越多地被用于证明就绪性和发布决策,然而它们往往未指明支持某项主张的证据是什么、该证据能确立什么,或该主张如何导向决策。我们提出EIO-Agents,一个基于两层架构的开放规范,用于可互操作的AI智能体评估。评估智能本体(EIO)通过类型化证据、版本化行为谓词、证据契约、主张、见证规则、证明状态、复发以及针对指标、发现、控制和PASS、REVIEW或BLOCK决策的可计算推导,提供语义层。可移植评估记录(PER)提供记录系统:一种规范化的、内容寻址的评估表示,保留从证据到决策的链条,并可重新推导、解释和验证。分数用于总结,陪审团用于解释,轨迹用于记录,但三者均未定义证据的含义或其能证明的内容。EIO提供了这一缺失的语义契约,而PER将产生的评估保留为可移植且可验证的记录系统。随着AI智能体承担更大的操作责任,评估必须超越分数和裁决的集合,成为可独立核查其含义、证据、局限性和决策的负责任工件。

英文摘要

AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used to justify readiness and release decisions, yet they often do not specify what evidence supports a claim, what that evidence can establish, or how the claim leads to a decision. We introduce EIO-Agents, an open specification for interoperable AI agent evaluation built on two layers. The Evaluation Intelligence Ontology (EIO) provides the semantic layer through typed evidence, versioned behavioral predicates, evidence contracts, claims, witness rules, proof status, recurrence, and computable derivations for metrics, findings, controls, and PASS, REVIEW, or BLOCK decisions. The Portable Evaluation Record (PER) provides the system of record: a canonical, content addressed representation of one evaluation that preserves the evidence to decision chain and can be re derived, explained, and verified. Scores summarize, juries interpret, and traces record, but none of them define what the evidence means or what it can prove. EIO provides that missing semantic contract, while PER preserves the resulting evaluation as a portable and verifiable system of record. As AI agents assume greater operational responsibility, evaluation must become more than a collection of scores and verdicts; it must become an accountable artifact whose meaning, evidence, limitations, and decisions can be independently checked.

Comments32 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑