arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34198cs.LG

冻结的裁判,移动的智能体:版本依赖的LLM裁判误差与裁判辅助智能体评估的局限性

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

Jiapeng Li

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示固定LLM裁判在评估智能体升级时存在版本依赖误差,提出采用显式参考标准和配对审计替代仅裁判决策,以提高评估可靠性。

中文摘要 AI 辅助

语言模型裁判将智能体升级版本与其前身进行比较,但固定的裁判可能会产生依赖版本的错误。我们分析了35个公开的编码智能体提交(20个预定义的版本对,涉及250个SWE-bench Verified问题)、两个客服智能体(155个tau-bench任务)以及1,106条专家标注的AgentRewardBench轨迹。一次上游中断导致用于主要SWE-bench分析的三个裁判(8,743个对齐的智能体-任务单元)受到影响;第四个裁判仅作描述性使用。所有三个编码智能体裁判和所有四个tau-bench裁判在多重性调整后均拒绝了任务条件误差不变性。在SWE-bench上,60个裁判-对单元中有32个具有可检测的差分比较成分;八个仅裁判区间宣称了执行区间无法确立的改进,尽管秩相关系数为0.71-0.79。在tau-bench中,一个裁判通过惩罚奖励忽略的程序性习惯,自信地逆转了九点的参考奖励差距。失败编码补丁的虚假接受率随智能体能力(以任务和参考结果为条件)而上升,而AgentRewardBench的任务可解性预测在SWE-bench中符号反转。将旧版本校准迁移到SWE-bench上使平均绝对比较误差从3.8个百分点上升到19.5个百分点;在修正边界附近,24.6%的比率自助抽样结果未定义。一个调整后的配对审计在80个标注任务上将经典区间仅收窄约5%。一个随机三臂测试不支持展示智能体最终报告会增加虚假接受率的预测(所有三个Holm调整p值均为1.0)。这些结果支持使用显式参考标准和当前输出的配对审计,而非仅依赖裁判的发布决策或迁移旧版本校准。

英文摘要

Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories. An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned cells); the fourth is descriptive. All three coding-agent judges and all four tau-bench judges reject task-conditioned error invariance after multiplicity adjustment. On SWE-bench, 32 of 60 judge-by-pair units have a detectable differential comparison component; eight judge-only intervals declare upgrades that execution-based intervals cannot establish, despite rank correlations of 0.71-0.79. In tau-bench, one judge reverses a nine-point reference-reward gap by penalizing a procedural habit the reward ignores. False acceptance of failed coding patches rises with agent capability conditional on task and execution outcome, while a task-solvability prediction reverses sign across domains. A separately fixed post-submission OpenHands follow-up on the same 250 issues (eight configurations, 1,981 three-judge cells) reproduces the capability/false-acceptance association (mean Spearman +0.944, exact p=0.000099) and decreasing Youden contrast (mean -0.937, p=0.000397); this is observational, not a new-task replication. Transporting old-version calibration raises SWE-bench comparison error from 3.8 to 19.5 points, with 24.6% undefined bootstrap ratios. A paired audit saves only 5% in interval width at 80 labeled tasks. A randomized self-report test is negative (three adjusted p-values=1.0). These results favor paired audits of current outputs over judge-only release decisions or transported calibration; independent human patch review remains pending.

补充信息

↑