发表机构
Columbia University; Johns Hopkins University; The University of Chicago; The Pennsylvania State University; The Hong Kong Polytechnic University(哥伦比亚大学; 约翰斯·霍普金斯大学; 芝加哥大学; 宾夕法尼亚州立大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出,删失评分量表上的双重差分法会制造虚假效应,通过预注册的LLM评判审计案例,发现名义显著的交互作用实为删失导致的偏差,推导了机制且其贡献可测量。
AI 中文摘要
LLM评判审计通过对比匹配条件来验证偏差,最强的设计采用双重差分:在同一项目内对比两个候选响应,再跨操纵属性进行差分,从有界评分量表中读取结果。我们表明,该终点在报告它的量表上是不可识别的。双重差分的每个项都受自身份额删失,因此观测到的统计量混淆了差异偏好与差异衰减:当两个响应受删失程度不等时,两者共有的严重程度偏移会产生交互作用,而这种不等性正是良好刺激所处的边界距离情况。我们在一项预注册的冻结教学法评判审计中展示了这一失效,该审计在990次调用的第一次前就已封存。注册的主要终点是陈述学习者画像对评判者支架偏好的效应,为0:+0.085分(95%BCa置信区间[-0.167, +0.353],p=0.684)。审计中一个名义上显著的交互作用+0.378(p=0.002)并非偏好:一个包含零差异偏好的结构仅从观测到的严重程度偏移和量表下限就可复现其79%至85%。我们以闭式推导了该机制,并表明其贡献可从审计自身的评分中测量。
英文摘要
LLM-judge audits assess bias by comparing ratings across matched conditions. Difference-in-differences designs compare two candidate responses within each item and then compare that contrast across a manipulated attribute. We show in closed form that this endpoint need not identify differential preference on the latent scale when ratings are bounded. A severity shift common to both responses produces an observed interaction whenever the scale bounds attenuate it unequally. We examine this problem in a pre-registered audit comprising 990 calls to a frozen pedagogy judge. The registered primary analysis did not detect an effect of the stated learner profile on scaffolding preference. The only nominally significant secondary interaction concerned productive struggle ($+0.378$ points; $p = 0.002$). A post-hoc construction with zero differential preference reproduced 79 to 85\% of this interaction using the observed severity shift and the scale floor alone. The observed interaction therefore does not identify differential preference on the latent scale.
CommentsAccepted at NewInML Workshop at NeurIPS 2026