arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体评估的可信层

A Trust Layer for Agent Evaluation

Mohammadreza Sediqin, Shivali Dalmia, Srinivasa Karthikeya Reddy Kovvuri, Abhishek Mukherji

arXiv 2610.07274首次发表:更新:

发表机构

Centific Research(森蒂菲克研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对基准分数不可信的问题,提出附加的事后可信层框架,通过四项确定性检查验证分数真实性,实验表明仅22.6%的通过分数可信。

AI 中文摘要

确定性基准分数表明智能体获得了分数,但并未表明该分数是否实至名归、是否诚实报告,或在第二次运行时是否仍然成立。我们引入了智能体评估的可信层(Trust Layer for Agent Evaluation),这是一个附加的事后框架,在每个记录的分数旁边报告该分数是否应被相信。它验证四个属性:结果是否得到基准自身评分逻辑的支持;通过的答案是否通过可追踪的计算获得;智能体的完成声明是否与实际发生的情况相符;以及结果在重复执行下是否稳定。前三个属性仅使用已保存的工件;第四个属性则重新运行智能体。模型判断仅在多数投票下标记证据;所有裁决遵循确定性规则,且绝不修改记录的分数。将该框架应用于智能体最后考试(Agents' Last Exam)中108个任务上的五种智能体配置,每个模型都显示出没有可追踪计算的通过运行(比率相差十倍)、被证实的虚假完成声明以及不稳定的结果:18-46%的任务在五次运行中未保持在同一个分数段。只有22.6%的记录的通过分数通过了全部四项检查(95%置信区间15.0-32.6,n=84)。测量智能体能够做什么与验证它确实做了什么是不同的问题,而当前的基准只解决了前者。

英文摘要

Deterministic benchmark scores show that an agent received credit, but not whether that credit was earned, reported honestly, or would hold on a second run. We introduce a Trust Layer for Agent Evaluation, an additive post-hoc framework that reports, beside each recorded score, whether it should be believed. It verifies four properties: whether the result is supported by the benchmark's own grading logic, whether a passing answer was earned through traceable computation, whether the agent's completion claim matches what occurred, and whether the result is stable under repeated execution. The first three use only saved artifacts; the fourth re-runs the agent. Model judgments only label evidence under majority voting; all verdicts follow deterministic rules and never modify the recorded score. Applied to five agent configurations on 108 tasks from Agents' Last Exam, every model shows passing runs with no traceable computation (at rates varying tenfold), confirmed false completion claims, and unstable results: 18-46% of tasks do not stay in one score band over five runs. Only 22.6% of recorded passes clear all four checks (95% CI 15.0-32.6, n=84). Measuring what an agent can do and verifying that it did it are different problems, and current benchmarks address only the first.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑