arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29709cs.LGcs.CLstat.ME

经典测量理论误导LLM评判者的三种方式

Three Ways Classical Test Theory Can Mislead About LLM Judges

Louis Yiven Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示经典测量理论三种统计量在LLM评判者评估中的误用,指出其无法单独归因于评判者,并提出四条报告准则以保持归因清晰。

中文摘要 AI 辅助

一个LLM评判者根据评分标准对一组回答进行评分,其信度系数为0.52。这究竟测量了什么?评判者评估已开始借用经典测量理论中的信度统计量,但通常未说明每个统计量所假设的测量设计。我们表明,三种广泛可移植的统计量对评判者与对测试的含义不同,因为评判者情境重新安排了这些设计所依赖的角色。首先,基于评分标准元素计算的内部一致性系数不包含评分者层面。当将一位评判者的测量误差率固定为4.72%时,随着题库围绕该误差率重新设计,KR-20仍可在0.01至0.68之间变化;而改变评判者误差会使系数产生相当幅度的变动,因此题目设计与评判者误差无法被分别识别,任何单一值都不能被解读为评判者的属性。其次,可靠性指数Φ(λ)是方差成分的比率,有时与之等同的分类概率在我们的题库上与其相差0.25-0.43,在底层模型完全成立的模拟数据上相差0.17-0.30。第三,Livingston-Lewis准确度以同一工具上考生自身的真实分数为基准,因此将其与外部金标准比较会将评判者不可靠性与标准无效性混为一谈。回顾三篇最相关的评判者评估论文,我们未发现这些错误的已发表实例,这使得该警示具有前瞻性。一个无法归因于评判者的系数仍会向下游传播至部署决策和披露文件中。因此,我们最后提出四条报告准则,以保持归因与数字相关联。

英文摘要

Evaluations that use a large language model (LLM) as a judge have begun to borrow reliability statistics from classical test theory and its extensions. We examine three such statistics that need one administration and no gold labels. None of them can isolate the judge, because one judge under one prompt supplies no variance component of its own. Claude Haiku 4.5 judged 210 constructed short answers against ten-element checklists. On the 180 with parsed verdicts, the Kuder-Richardson coefficient (KR-20) came out at 0.5223 on the judge's verdicts and 0.5231 on error-free gold verdicts. In simulation, bank design alone moves KR-20 from 0.01 to 0.68 at the judge's measured 4.72% error rate. The dependability index $Φ(λ)$, a ratio of mean squared distances from the pass mark, sits 0.22 to 0.38 below the judge's accuracy against gold and returns 0.54 to 0.68 on error-free gold verdicts. Livingston-Lewis accuracy treats the rubric elements as a sample, and at a pass mark of five elements it credits error-free gold scores with 0.78, close to the judge's 0.81. A statement about the judge therefore needs gold labels or a varied scorer facet, and a reliability ratio needs the bank's spread beside it. One of the four closest judge-evaluation papers varies the prompt and still reads a reliability below 0.7 as a sign that a model cannot serve as a judge, although that reliability moves with the spread of the samples scored. We derive a decision table and four reporting lines from these two rules.

发表机构

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑