arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估超越分数对齐的LLM作为评判者:残差评判难度的心理测量学分析

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne

arXiv 2610.02877首次发表:更新:

发表机构

DIPF | Leibniz Institute for Research and Information in Education; Centre for International Student Assessment (ZIB); Faculty of Computer Science, Goethe University Frankfurt; Chemnitz University of Technology(DIPF | 莱布尼茨教育与研究信息研究所; 国际学生评估中心(ZIB); 法兰克福大学计算机科学学院; 开姆尼茨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文从心理测量学角度分析LLM作为评判者的可靠性,发现人类与LLM在残差评判难度上存在维度依赖的差异,提出残差诊断以补充总体对齐评估,支持更有效的人机协作。

AI 中文摘要

大型语言模型(LLMs)被广泛用作自动评判者,其有效性通常通过与人类评分的对齐程度来评估。然而,总体一致性无法揭示人类和LLM是否认为相同的评估案例存在难度。在本文中,我们从心理测量学的角度研究摘要评估中的这一问题。我们分别对人类和LLM的评分拟合多面Rasch模型,将分数分解为潜在摘要质量、评判者严格度、维度严格度和评分量表阈值。基于这一分解,我们将残差难度定义为一种模型调整后的评判难度度量,并比较人类与LLM评判者是否共享相同的难度结构。在SummEval数据集上的17个开放权重LLM评判者中,我们发现潜在摘要质量的中等对齐并不意味残差难度的对齐。人类和LLM评判者在哪些摘要-维度单元仍然困难方面存在差异,且这种不匹配强烈依赖于维度。一致性显示出显著的LLM困难偏移,而连贯性则显示出人类困难偏移。我们进一步表明,人类容易但LLM困难的案例可以从可观察的源-摘要属性中部分预测。这些发现表明,总体人类对齐仅反映了LLM作为评判者可靠性的一部分,而心理测量学残差诊断支持更具信息量的评判者评估和更有针对性的人机协作。

英文摘要

Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psychometric perspective. We fit Many-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating-scale thresholds. Building on this decomposition, we define residual hardness as a model-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure. Across 17 open-weight LLM judges on SummEval, we find that moderate alignment in latent summary quality does not imply alignment in residual hardness. Human and LLM judges differ in which summary--dimension units remain difficult, and this mismatch is strongly dimension-dependent. Consistency shows a pronounced LLM-hard shift, whereas coherence shows a human-hard shift. We further show that human-easy but LLM-hard cases are partially predictable from observable source--summary properties. These findings suggest that aggregate human alignment reflects only part of LLM-as-a-judge reliability, while psychometric residual diagnostics support more informative judge evaluation and more targeted human--LLM collaboration.

CommentsAccepted at AACL-IJCNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑