作为评分者的大语言模型评判者:针对公开语料库上大语言模型作文评分的严苛性、光环效应、可靠性及版本不稳定性的预注册审计
LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora
浏览论文内容
中文总结 AI 辅助
该研究将LLM视为评分者,通过预注册审计发现LLM作文评分存在严苛性、版本不稳定性等问题,其与人类评分者相关性低,且未达人类水平准确性。
中文摘要 AI 辅助
大语言模型(LLM)正越来越多地被用作学习分析中的作文评分者,其评估几乎仅采用一致性统计。教育测量领域提醒我们,评分者还存在严苛性差异、光环效应以及作为工具的漂移现象。我们将LLM评判者视为评分者,并在两种语言的公开语料库(ENEM/Essay-BR;ASAP)上开展预注册的评分者效应系列分析(多 facet Rasch 严苛性分析、残差光环效应、概化理论/决策研究、跨版本偏移、差异功能):共2377篇作文、12名评判者、4家提供商、5个版本对比、重复单元,作为评分张量发布。在ENEM的0-1000分制下,评判者严苛性跨度达219分;在ASAP上,评判者群体的严苛性跨度为分数范围的15%-33%,而受过训练的人类评分者间差距接近1%。评判者与人类的相关性处于0.47-0.56的无区分度区间。全部5项版本对比的严苛性偏移均超出家族式置换零假设(最高达133分),且1名评判者在研究中期被弃用,通过身份金丝雀检测到。两项预注册测试返回了真实的零结果:经严苛性调整的排行榜反转未通过置换零假设检验,且“无声漂移”被反驳:5项对比中有4项的一致性随严苛性变化。重复实验产生了自我一致性(k≤2时φ≥0.80),但未达到人类水平的准确性,且同工具检查推翻了我们自身的光环效应对比:在工具和校准匹配的情况下,我们未发现可信证据表明评判者的光环效应超过受过训练的人类范围。
英文摘要
Large language models (LLMs) are increasingly used as essay graders in learning analytics, evaluated almost exclusively with agreement statistics. Educational measurement warns that raters also differ in severity, show halo, and drift as instruments. We treat LLM judges as raters and run a pre-registered rater-effects battery (many-facet Rasch severity, residual halo, generalizability/decision studies, cross-version shifts, differential functioning) on public corpora in two languages (ENEM/Essay-BR; ASAP): 2,377 essays, 12 judges, 4 providers, 5 version contrasts, replicated cells, released as a score tensor. Judge severity spans 219 points on ENEM's 0-1000 scale; on ASAP the panel spread is 15-33% of the score range against a between-trained-human gap near 1%. Judge-human correlations sit in an undiscriminating .47-.56 band. All five version contrasts shift severity beyond a family-wise permutation null (up to 133 points), and one judge was deprecated mid-study, caught by identity canaries. Two pre-registered tests returned honest nulls: severity-adjusted leaderboard reversals did not survive a permutation null, and "silent drift" was refuted: agreement moved with severity in four of five contrasts. Replication yields self-consistency (phi>=.80 at k<=2) but not human-level accuracy, and a same-instrument check overturned our own halo comparison: matched on instrument and calibration, we find no credible evidence that judge halo exceeds the trained-human range.