arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

trajectory-judge:仅基于结果的LLM评估器在智能体轨迹上的缺失

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

Hadi Mohammadi

arXiv 2609.00038首次发表:更新:

发表机构

Utrecht University(乌得勒支大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

trajectory-judge研究发现仅基于结果的LLM评估器对智能体轨迹的静默故障捕获率低,提出的步骤rubric评估器能以零误报实现更高的静默召回,同时发布相关环境与分析工具。

AI 中文摘要

仅基于结果的评估是LLM智能体的生产默认方案:向评估器展示请求和最终回复,询问是否处理得当。该指标从结构上无法识别那些以错误方式得出正确答案的智能体。我们利用构造已知真实值的环境来测量这一盲区:确定性的使用工具的客服环境、始终能解决问题的脚本式神谕策略,以及在已知步骤恰好破坏一个元素的故障注入器,根据客户可见结果是否保留(静默故障)或不保留(明显故障)对故障进行分层。我们在400条轨迹上对五种评估器(程序规则、仅基于结果、两种模型规模的步骤 rubric、自一致性集成)在检测、步骤定位、故障类型、校准和成本方面进行评分:仅基于结果的评估器能捕获84%的明显故障,但仅捕获45%的静默故障,同时标记了33%的正确轨迹;步骤 rubric 评估器实现了77%的静默召回,且零误报,成本是前者的3倍。没有评估器读取最终回复:在原本完美的轨迹中添加的虚假承诺完全规避了规则和步骤评估器,占比达82%,而自一致性使成本增加了两倍却无任何提升。我们认为评估器评估必须按结果保留情况对召回率进行分层,并发布该环境、注入器、所有原始裁决以及可离线重建所有数值的分析流程。

英文摘要

A direct test of an LLM judge of agent trajectories injects faults into correct runs and reports recall, per fault type or by whether the fault broke the environment outcome (loud) or not (silent). Such recall can credit a judge with detection it does not have; paired discrimination, its flag rate on the faults minus its rate on the clean runs they came from, exposes this. Our testbed, a deterministic support desk with a scripted oracle and a one-step fault injector, labels all 400 trajectories exactly. A 14B judge shown only the request and final reply scores 34% to 76% recall on four fault types that leave the reply unchanged. There its input is the clean run's, so its paired discrimination is zero and that recall is its flag rate on clean runs. Splitting by outcome survival does not fix this: its loud recall of 84% is a paired +0.393 and its silent recall of 45% a paired +0.048, all from the two fault types that change the reply. Told to check each step, the same model flags every fault of those four types and 0 of 100 clean runs (95% CI up to 3.6%). It does not reliably check the reply: of four invented promises it flags one every time and the other three once in 42 faults. Shown every step but asked only about the reply, it still reaches a paired +0.69 on reply-unchanged faults, against +1.00 when told to check each step. We recommend reporting paired discrimination against clean parents, split by whether the fault reaches the judge's input and by outcome survival, and release the testbed, raw verdicts and analysis pipeline.

CommentsAccepted at the NeurIPS 2026 Workshop: Who Verifies the Agents? Toward Reliable Agent Development (poster). Camera-ready version. 22 pages, 5 figures, 14 tables. Code and data: https://github.com/mohammadi-hadi/trajectory-judge

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑