JEV-as-a-Judge:自信时接受,不确定时升级
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
浏览论文内容
中文总结 AI 辅助
本文提出JEV仅决策评判者,以极低成本接近最先进LLM评判者,并通过置信度级联升级不确定裁决,保留99%准确率。
中文摘要 AI 辅助
LLM-as-a-judge(大语言模型作为评判者)能够在多种任务中实现评估,但在大规模应用中,推理成本和置信度可靠性变得至关重要。我们研究了一个仅做决策的评判者(decision-only judge)能否提供经济高效的初步筛选,并识别出何时需要更强的评估。通过将jev-as-a-judge与十六种生成式评判者和奖励模型评判者进行比较,并采用盲法人工仲裁,我们发现,在普通偏好和基于证据的事实性任务上,其表现与最先进的LLM评判者(我们最强的比较对象)相差不到三个百分点,而费用仅为该比较对象的0.36%。当判断需要检查推导过程或抵御精心编写的错误答案时,差距会更大。在多个基准测试中,JEV与该比较对象之间的差距集中在低置信度决策上。一个冻结的级联模型(frozen cascade)接受高置信度的裁决,并将不确定的裁决升级处理,在较低成本下保留了该比较对象99%的准确率。
英文摘要
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。