arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13564cs.AI

生成无奖励的评判标准以减少智能体评估中的过度打分

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出 RubricForge 方法,从带标签轨迹生成无奖励的智能体评判标准,可减少智能体评估中的过度打分,在 tau-bench 和 WebShop 上的表现优于通用 G-Eval 评判器。

中文摘要 AI 辅助

大规模评估语言模型智能体越来越依赖第二种语言模型作为自动评判器,因为黄金信号(即可执行的环境奖励)在部署时成本高、速度慢或不可用。这种评判器是一种无奖励代理,其价值取决于是否可信,但现有的评判器要么像 G-Eval 那样手动编写评分标准,要么微调评判器的权重,两者都倾向于将流畅但未成功的轨迹评为成功。我们转而从一小部分带有真实标签的轨迹中生成智能体评判标准的文本,使其基于真实结果。我们提出了 RubricForge,它通过对带标签的轨迹进行反思性演化来生成评判标准,以最大化与环境奖励的一致性,随后冻结该标准,并在不访问环境的情况下,通过一次模型调用将其应用于保留的轨迹。优化后的产物是人类可读的文本,因此每个评判都可归因于明确的标准。使用一个冻结的 7B 模型同时作为智能体和评判器,在 tau-bench(来自 220 次 rollout 的 173 个带标签轨迹)和 WebShop(160 个)上,主要提升在于忠实度而非原始一致性。与通用 G-Eval 评判器相比,其优势在统计上不显著(McNemar 检验 p=0.248),绝对分数校准略微偏向通用评判器(|err| 差值 -0.048,p=2×10^-4)。然而,RubricForge 将失败轨迹过度打分的频率约为一半(tau-bench 上的误通过率为 0.115 对比 0.173,有 3 次过度打分捕获,0 次反转),且能更忠实地对 WebShop 的分级结果进行排序(Spearman 相关系数 0.410 对比 0.370)。对于无奖励评估器而言,误通过率而非总体一致性是与部署相关的关键指标,因为误通过会交付有缺陷的智能体,而误不通过仅会导致一次重试。

英文摘要

Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time. Such a judge is a reward-free proxy whose value depends on whether it can be trusted, yet existing judges either hand-write the scoring rubric, as in G-Eval, or fine-tune the judge's weights, and both tend to credit fluent but unsuccessful trajectories as successes. We instead induce the text of an agent-judging rubric from a small set of ground-truth-labeled trajectories, grounding it in true outcomes. We present RubricForge, which evolves a judge rubric by reflective evolution against labeled trajectories to maximize agreement with the environment reward, freezes it, and applies it to held-out trajectories in one model call with no environment access. The optimized artifact is human-readable text, so every verdict is attributable to named criteria. Using one frozen 7B model as both agent and judge, on tau-bench (173 labeled trajectories drawn from 220 rollouts) and WebShop (160), the principal gain is faithfulness rather than raw agreement. The edge over a generic G-Eval judge is not statistically significant (McNemar p = 0.248), and absolute-score calibration marginally favors the generic judge (|err| difference -0.048, p = 2x10^-4). Yet RubricForge over-credits failed trajectories roughly half as often (0.115 vs. 0.173 false-pass rate on tau-bench, with three over-credit catches and zero reversals) and ranks graded WebShop outcomes more faithfully (Spearman 0.410 vs. 0.370). For a reward-free evaluator the false-pass rate, not aggregate agreement, is the deployment-relevant quantity, since a false pass ships a broken agent whereas a false fail merely costs a retry.

发表机构

  • Trinity College Dublin(都柏林三一学院)
  • University College Dublin(都柏林大学学院)
  • Dublin City University(都柏林城市大学)

机构由 AI 辅助整理,请以论文原文为准。

↑