选择塑造边界:未选自然语言推理群体中单调性和标签一致性的预注册复制
Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations
浏览论文内容
中文总结 AI 辅助
研究自然语言推理中人类标签变异,预注册复制未选群体中单调性和标签一致性边界,原负边界预测失败,新结果显示为正,表明原负边界或为低一致性选择结构,相关主张应明确选择条件。
中文摘要 AI 辅助
先前关于自然语言推理(NLI)中人类标签变异(HLV)的研究常常依赖于通过分歧程度选择项目的重新注释资源。一项早期研究发现,包含非向上单调算子的假设在ChaosNLI(Cliff's delta = -0.284)中显示出较低的标签一致性,ChaosNLI仅限于多数标签恰好获得五票中的三票的项目。我们预注册了在ChaosNLI所源自的未选群体(SNLI和MultiNLI开发集)中对这一边界的复制,使用相同的算子标记器和四级有序一致性结果。注册的预测失败了。所有七个对比都返回正的Cliff's delta(非向上项目的一致性略高,而非更低),唯一显著的验证性对比与注册的符号相反,并且每个效应都远低于我们感兴趣的最小效应大小(0.10)。稳健性检查支持该测量:模拟标记器错误分类会缩小效应而非制造效应,并且在一个新的200项样本上手动重新标记审核达到了0.875的四类一致性。我们得出结论,早期的负边界可能是基于低一致性选择的结构,而非群体层面的属性,并且基于所选重新注释资源构建的HLV结构主张应明确说明其选择条件。
英文摘要
Prior work on human label variation (HLV) in natural language inference (NLI) has often relied on re-annotation resources that select items by disagreement level. An earlier study (arXiv:2607.15870) found that hypotheses containing non-upward monotonicity operators showed lower label agreement in ChaosNLI (Cliff's delta = -0.284), which is restricted to items whose majority label carries exactly three of five votes. We preregistered a replication of this boundary in the unselected populations that ChaosNLI was drawn from: the SNLI and MultiNLI development sets, using the same operator tagger and a four-level ordinal agreement outcome. The registered prediction fails. All seven contrasts return a positive Cliff's delta (non-upward items agree slightly more, not less), the only significant confirmatory contrast has the opposite sign to the registration, and every effect is far below our smallest effect size of interest (0.10). Robustness checks support the measurement: simulated tagger misclassification shrinks the effects rather than manufacturing them, and a manual re-tagging audit reaches four-class agreement of 0.875 on a fresh 200-item sample. We conclude that the earlier negative boundary is plausibly a structure conditional on low-agreement selection rather than a population-level property, and that HLV structure claims built on selected re-annotation resources should state their selection conditional explicitly.
发表机构
- University of Bremen(不来梅大学)
机构由 AI 辅助整理,请以论文原文为准。