arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

条件准确率画像:诊断不同部署条件下的LLM裁判

Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions

Wenqi Li, Bin Liu, Mindi Ruan, Chuanbo Hu, Minglei Yin, Xin Li

arXiv 2610.09229首次发表:更新:

发表机构

University at Albany, State University of New York; West Virginia University(纽约州立大学奥尔巴尼分校; 西弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出条件准确率画像(CAP)框架,将LLM裁判准确率分解为八种部署条件,揭示聚合准确率隐藏的差异,为选择裁判提供更可操作的依据。

AI 中文摘要

LLM作为裁判现已成为可扩展评估的标准工具,但裁判性能仍常被概括为单一的准确率数值。这种聚合视图掩盖了裁判成功或失败所依赖的部署条件。我们提出条件准确率画像(CAP),这是一种事后诊断框架,将成对LLM裁判准确率分解为八种条件,这些条件组织为内容敏感性、鲁棒性和理由质量三个方面。CAP与基准无关:当基准提供所需标注时可直接应用,可通过任务子集代理近似应用,或在可生成扰动对时通过受控增强应用。我们在七个LLM裁判和六个成对评判基准上实例化CAP,包括我们创建的用于支持全部八种条件的受控测试平台judgerEva-Standard。CAP揭示了被聚合准确率隐藏的画像差异:在judgerEva的裁判无关Hard-Constructed子集上,对遗漏限定条件最敏感的两个裁判在七个裁判中按总体准确率排名垫底三位,因此遗漏敏感性无法由聚合准确率预测。跨基准而言,位置鲁棒性显示出最强的排名稳定性(平均Spearman ρ̄=0.87),但在JudgeBench-Pro对抗压力下本身脆弱,在共享条件中显示出最大的平均准确率下降,尽管主要退化通道因裁判而异。条件级画像为选择LLM裁判提供了比聚合准确率更具可操作性的依据。

英文摘要

LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number. This aggregate view hides the deployment conditions under which a judge succeeds or fails. We introduce \textbf{Conditional Accuracy Profiling} (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality. CAP is benchmark-agnostic: it can be applied directly when a benchmark provides the required annotations, approximately through task-subset proxies, or through controlled augmentation when perturbation pairs can be generated. We instantiate CAP on seven LLM judges across six pairwise judging benchmarks, including \textsc{judgerEva-Standard}, a controlled testbed we created to support all eight conditions. CAP exposes profile differences hidden by aggregate accuracy: on \textsc{judgerEva}'s judge-independent Hard-Constructed subset, the two judges most sensitive to omitted qualifications rank in the bottom three of seven by overall accuracy, so omission sensitivity is not predicted by aggregate accuracy. Across benchmarks, Position Robustness shows the strongest rank stability (mean Spearman $\barρ{=}0.87$) but is itself fragile under JudgeBench-Pro adversarial stress, showing the largest mean accuracy drop among the shared conditions, though the dominant degradation channel varies by judge. Condition-level profiles provide a more actionable basis than aggregate accuracy for selecting LLM judges.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑