arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27463cs.AI

评估评估者:用于大语言模型评估的拉施测量理论

Rating the Raters: Rasch Measurement Theory for LLM Evaluation

Pratik S. Sachdeva, Nathan Boudol

首次发表
浏览论文内容

中文总结 AI 辅助

该研究将拉施测量理论应用于LLM评估,通过多维度拉施模型分析发现LLM与人类评估者存在系统性差异,提出RMT适用于各类LLM评估范式。

中文摘要 AI 辅助

大语言模型(LLM)如今参与到评估的各个环节:作为在基准测试中被评分的考生、作为其他模型输出的评判者,以及作为人类生成内容的评估者。每种范式都可被视为一个测量问题,其中评估者借助工具(如基准测试)中的条目来探测对象的潜在属性。标准评估实践往往忽略每个核心组件对最终结果的贡献,限制了我们对测量对象的理解。拉施测量理论(RMT)非常适合这类问题,它将有序评分分解为共同量表上的可分离维度,还提供了一系列诊断方法,可识别校准不当的测量和评估者偏差。我们开展了一项案例研究,将RMT应用于“LLM作为评估者”的范式,使用的数据集为Measuring Hate Speech语料库,其构建本身就基于RMT。我们对来自9个涵盖不同家族和能力水平的LLM的标注结果,拟合了一系列多维度拉施模型。分析显示,LLM在严格程度、条目级校准、问题顺序鲁棒性、目标身份敏感性以及评分量表使用方面,均与人类评估者存在系统性差异,而这些差异会被标准评估实践所掩盖。总体而言,我们认为RMT应纳入评估“LLM作为考生、评判者及评估者”范式的工具集。

英文摘要

LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by judges or raters. Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured. Rasch measurement theory (RMT) is well-suited to this problem. RMT decomposes ordinal ratings into separable facets on a common scale. It further provides a battery of diagnostics that can identify miscalibrated measurements and rater biases. We present a case study of RMT applied to the LLM-as-rater paradigm using the Measuring Hate Speech corpus, whose construct was itself built under RMT. We fit a series of many-facet Rasch models to annotations from nine LLMs spanning families and capability levels. Our analyses show that LLMs systematically differ from human raters in severity, item-level calibration, question-order robustness, target-identity sensitivity, and rating scale use, all of which standard evaluation practice would largely obscure. Overall, we argue that RMT belongs in the toolkit for evaluating LLM-as-examinee, -judge, and -rater paradigms.

发表机构

  • D-Lab, University of California, Berkeley(加州大学伯克利分校D-Lab)
  • Grenoble INP(格勒诺布尔国立综合理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑