TRACE:诊断智能体评估中验证器的脆弱性
TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation
- Prime Intellect
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出TRACE协议,用于诊断智能体评估中验证器的脆弱性,通过针对性修改评估部分内容,可区分评分变化源于智能体还是测量本身,能检测真实效果。
AI中文摘要:
如今,验证器评分既作为大型语言模型(LLM)智能体的基准指标,也作为训练奖励,评分的变化常被解读为能力的变化,但实际上可能反映的是评估方式的变化。本文提出TRACE协议,该协议将评分变化从一种判定转化为可检验的诊断:它对评估的某一部分进行针对性修改,对比配对运行结果,检查智能体的行为是否发生变化,并对未改变的轨迹重新评分以检验评分规则是否是导致变化的原因。在包含25个合成任务的受控套件中,重命名工具会使脚本化智能体的评分降低0.250,尽管其执行的操作完全相同;在评分时恢复原始名称可完全消除该差距,而相同的修改会暴露第二个智能体的真实行为失败。在包含4个LLM智能体的公开τ²-bench任务中,一项初始30任务研究发现了混合奖励变化,其中唯一明确的效果未被复现。在针对88个新任务的更大规模后续研究中,每个条件下重复运行,重命名工具或重新格式化工具输出对8组智能体-条件对中的7组的奖励影响在±0.10以内,而故意误导的工具名称会使每个智能体的奖励降低0.20-0.44,表明该设置可检测真实效果。相同的重复运行会使15-36%的任务结果翻转,因此单次运行比较无法区分呈现效果与运行间变化。两个前沿LLM评判者在呈现固定轨迹不同时给出一致判定,但在57%的相同记录上彼此意见不一致,主要原因是其中一个评判者按程序而非结果评分。因此,TRACE可区分评分变化所反映的智能体情况与测量情况。
英文摘要:
Verifier scores now serve as both benchmark metrics and training rewards for large language model (LLM) agents, and a change in score is routinely read as a change in capability. It may instead reflect a change in the evaluation. We introduce TRACE, a protocol that turns a score change from a verdict into a testable diagnosis: it applies a targeted change to one part of an evaluation, compares paired runs, checks whether the agent's behavior changed, and rescores unchanged trajectories to test whether the scoring rule is responsible. In a controlled suite of 25 synthetic tasks, renaming tools lowers a scripted agent's score by 0.250 even though it performs exactly the same operations; restoring the original names at scoring time closes the entire gap, while the same mutation exposes a genuine behavioral failure in a second agent. On public $τ^2$-bench tasks with four LLM agents, an initial 30-task study finds mixed reward changes whose one clear effect does not replicate. In a larger follow-up on 88 new tasks with repeated runs per condition, renaming tools or reformatting tool outputs leaves reward unchanged to within $\pm$0.10 for seven of eight agent-change pairs, whereas tool names that deliberately mislead lower every agent's reward by 0.20-0.44, showing that the setup can detect real effects. Identical reruns flip 15-36% of task outcomes, so single-run comparisons cannot separate presentation effects from run-to-run variation. Two frontier LLM judges give consistent verdicts when a fixed trajectory is presented differently, yet disagree with each other on 57% of the same records, largely because one grades procedure rather than outcome. TRACE thus separates what a score change says about the agent from what it says about the measurement.