arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24419cs.AI

评审应知晓变化内容:LLM作为评审评估的结构效度

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM作为评审的评估,提出二维结构效度指标(不变性S与结构敏感性R),发现高评审一致性常伴随对结构变化的弱敏感性,建议联合报告两指标并审计验证集。

中文摘要 AI 辅助

LLM作为评审(LLM-as-a-judge)的评估通常通过一致性和对表层扰动的鲁棒性来衡量,但可靠性并不能确立结构效度。我们将评估者的结构效度形式化为二维轮廓:不变性S,即在保留结构的编辑下裁决保持不变的概率;以及结构敏感性R,即在最小化改变结构的编辑下裁决发生变化的概率。我们证明S和R相互独立,且没有标量摘要能保留所有相关比较。我们使用7种改变结构的干预类型和5种仅调整语域的对照,在7个评审和4个领域中测量该轮廓,干预方向由人工标注者确定,生成、验证和评审任务分配给不重叠的模型家族。在匹配的不变性S≥0.90时,评审的平均S为0.945,但R为0.319。敏感性在范围编辑和强度编辑之间也存在差异:R范围=0.383,而R强度=0.262,7个评审均呈现+0.121的相同符号差距。我们进一步审计了5个公开标签集,发现仅表层的预测器在配对模式下可复现55%-67%的标签,包括MT-Bench的67.4%人工投票。这些结果表明,高评审一致性可与对被评估结构变化的弱敏感性共存,这促使人们联合报告不变性和敏感性,并对验证集本身进行审计。

英文摘要

LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile: invariance S, the probability that a verdict is unchanged under construct-preserving edits, and construct sensitivity R, the probability that it changes under minimal construct-changing edits. We show that S and R are independent and that no scalar summary preserves all relevant comparisons. We measure the profile across 7 judges and 4 domains using 7 construct-changing intervention types and 5 register-only controls, with intervention direction determined by human annotators and generation, verification, and judging assigned to disjoint model families. At matched invariance S >= 0.90, judges average S = 0.945 but R = 0.319. Sensitivity also differs between scope and strength edits: R_scope = 0.383 versus R_strength = 0.262, a +0.121 gap with the same sign for all 7 judges. We further audit five public label sets and find that surface-only predictors reproduce 55%-67% of labels in paired mode, including 67.4% of MT-Bench human votes. These results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.

发表机构

  • South China University of Technology(华南理工大学)
  • University of Macau(澳门大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑