arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29769cs.CL

JEV vs. LLMs 作为评分标准评审:更便宜、更快,且在相同位置出错

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Delip Rao, Chris Callison-Burch

AI总结:

本研究比较 Jev 分类器与 LLM 评审在评分标准上的表现,发现 Jev 更便宜、更快且准确率相当,但相关错误限制了级联优势,仅能小幅提升。

AI中文摘要:

我们探讨 Jev(一种类型化分类器,它返回允许答案上的概率而不生成文本)能否替代 LLM 评分标准评审。我们将它与三个 flash 级 LLM 评审在来自七个基准的九个面板上进行比较,给每个评审相同的标准文本。在 27 个配对比较中,Jev 的准确率仅在 8 个上与 LLM 评审有显著差异,主要在二元标准上领先,仅在分级标准上落后,而大多数其他比较结果不具决定性。在九个面板上汇总,LLM 评审每个标准调用一次,成本是 Jev 的 29 到 325 倍,耗时是 Jev 的 30 到 220 倍。在分级标准上,所有四个评审彼此之间的一致性高于与标签的一致性,并且大多分配比评分者更低的等级。几种观察性解释之一认为,评分者遵循了我们标准文本中未包含的评分惯例。Jev 的置信度在大多数面板上对其自身错误进行排序,这应使廉价分类器成为级联的理想第一阶段,该级联将其不确定的裁决推迟给 LLM 评审。相关错误抵消了这一优势。LLM 评审几乎重复了 Jev 最自信的错误,因此在记录裁决上重放的级联降低了成本,但在交叉拟合阈值下最多比最佳单一评审提高 1.5 分,即使使用 oracle 阈值也最多提高 2.0 分。

英文摘要:

LLM judges score outputs against rubrics well enough to have become the norm, both in benchmarks and as rewards for training. Jev, a classifier-like alternative its creators call a "decision model", returns probabilities over permitted answers with a calibrated confidence score, which LLM judges do not natively provide. We compare Jev with three flash-tier LLM judges on nine panels from seven benchmarks with human judgments, giving every judge identical criterion texts. The LLM judges run in two setups: holistically, reading a whole rubric at once as Jev does, and one criterion at a time. Jev can often stand in for them. They cost 16 to 325 times as much and take 28 to 350 times as long, yet in each setup Jev's accuracy differs significantly from theirs in at most 8 of 27 paired comparisons, ahead mostly on binary checklist criteria and behind only on ordinal ones. Despite their different designs, the two kinds of judge err alike. On ordinal criteria, all LLM judges and Jev depart from the human raters together, agreeing more with one another than with the labels and mostly assigning lower levels. On Jev's most confident errors, about 96% of LLM verdicts repeat its wrong answer, where independent errors would give about half. Intuitively, calibrated confidence should make Jev an ideal first stage of a cascade that defers uncertain verdicts to an LLM judge. Yet such cascades only lower cost while adding little accuracy: even with oracle thresholds, none beats the best single judge by more than 2.7 points. Calibration can tell a cascade when to defer, but the cascade also needs a fallback that errs elsewhere; these judges are wrong in the same places. These findings, which hold in both setups and at high reasoning effort, suggest that a cascade of judges succeeds only when its judges make complementary errors, and that future decision models should be designed afresh with that aim.

补充信息

↑