发表机构
University of California, Los Angeles; Tsinghua University; LA General Hospital; Google(加州大学洛杉矶分校; 清华大学; 洛杉矶总医院; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过短篇故事心理深度评估发现,LLM评判者偏向推理输出但忽略人类主观异质性,开发集性能不足以证明部署有效性。
AI 中文摘要
LLM-as-a-Judge评估器越来越多地被用于对开放式生成进行评分,然而,当输出高度匹配且人类偏好具有主观性时,评判者在其开发集上与人评分的相关性可能无法保证测量的有效性。我们通过短篇故事中的心理深度来研究这种失效模式。七位人类读者和一个在原始心理深度量表数据集上选出的LLM评判集成(ρ=0.646)评估了来自GPT-5与GPT-4o以及DeepSeek-R1与DeepSeek-V3的60对盲法、提示匹配的故事对。人类偏好未显示出普遍的推理优势:GPT-5略优于GPT-4o(60.0%–62.9%),而DeepSeek-R1落后于V3(42.9%),且读者间一致性接近随机水平(Krippendorff's α=0.070),读者内部的一致性和重复出现的权重模式表明存在结构化异质性而非随机应答。相比之下,评判者在89.0%的维度级比较中偏向推理输出,在总体PDS上偏向60对中的59对,在所有五种评估配置中均一致,且其分数与句子长度和词汇多样性等表面特征相关。这些结果表明,开发集性能不足以证明在偏移分布上的部署有效性,且点估计评判者可能掩盖主观人类评估中的异质性。
英文摘要
LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short stories. Seven human readers and an LLM-judge ensemble selected on the original scalar Psychological Depth Scale dataset ($ρ= 0.646$) evaluated 60 blinded, prompt-matched story pairs from GPT-5 vs.\ GPT-4o and DeepSeek-R1 vs.\ DeepSeek-V3. Human preferences showed no universal reasoning advantage: GPT-5 was modestly preferred over GPT-4o (60.0--62.9\%), whereas DeepSeek-R1 trailed V3 (42.9\%), and inter-reader agreement was near chance (Krippendorff's $α= 0.070$), with within-reader consistency and recurring weighting patterns suggesting structured heterogeneity rather than random responding. The judge, by contrast, favored reasoning outputs in 89.0\% of dimension-level comparisons and 59 of 60 pairs on aggregate PDS, uniformly across all five evaluator configurations, and its scores were associated with surface features such as sentence length and lexical diversity. These results suggest that development-set performance is insufficient evidence for deployment validity on a shifted distribution, and that point-estimate judges can obscure the heterogeneity in subjective human evaluation.
Comments24 pages