arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JudgeProfile:理解并引导LLM评判者的主观性

JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, Yuanpeng Tu, Daniel Li, Junbiao Tang, Pengtao Xie, Zihao He

arXiv 2609.36705首次发表:更新:

发表机构

Meta; University of California San Diego; The University of Hong Kong; The Hong Kong University of Science and Technology(Meta; 加利福尼亚大学圣迭戈分校; 香港大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出JudgeProfile框架,将LLM评判分解为感知与优先级排序,通过SubjectiveSet数据集和重加权方法,将评判者决策与参考标准的一致性从66.48%提升至71.97%,优于微调和规则提示。

AI 中文摘要

LLM评判者本质上是主观的,在成对比较中,当两个选项都没有客观错误时,它们往往偏爱不同的回答。为了研究这种主观性,我们引入了JudgeProfile,一个将LLM评估分解为感知(评判者如何比较两个回答在清晰度、正确性和细节等特定属性上的表现)和优先级排序(每个属性对最终选择的影响程度)的框架。我们策划了SubjectiveSet,一个包含来自17个公共数据源的50,013个回答对的数据集,由21个LLM评判者依据87个属性进行评估。我们发现感知中存在一种隐藏的共识:即使评判者的整体选择存在分歧,他们也经常在属性判断上达成一致。基于这种分离,我们首先通过从评判者自身的整体选择中估计的属性权重来刻画每个评判者的优先级排序。即使从相同的属性判断中估计,这些权重在不同评判者之间也有所不同。然后,我们从参考标签中学习新的权重,以使其决策适应目标评估标准。重新加权感知属性将平均保留集上对参考标签的一致性从66.48%提高到71.97%,优于微调和基于规则提示的方法。我们的研究结果表明,理解和引导LLM评判者的主观性不仅需要关注它们感知到什么,还需要关注它们如何对其进行优先级排序。

英文摘要

LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87 attributes. We find a hidden consensus in perception: judges frequently agree on attribute judgments even when their overall choices diverge. Building on this separation, we first characterize each judge's prioritization using attribute weights estimated from its own overall choices. These weights differ across judges even when estimated from the same attribute judgments. We then learn new weights from reference labels to adapt their decisions to a target evaluation standard. Reweighting perceived attributes improves average held-out agreement with reference labels from 66.48% to 71.97%, outperforming fine-tuning and rubric prompting. Our findings show that understanding and steering the subjectivity of LLM judges requires attention not only to what they perceive, but also to how they prioritize it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑