AI 中文总结
该研究以桌游抽象任务为对象,发现多数语言模型虽具项目敏感性却与随机选择无显著差异,提出“无对齐的一致性”概念,指出相关评估存在缺陷,且字面相似度基线表现优于多数模型。
AI 中文摘要
项目敏感性(指模型的选择是否依赖于特定输入而非自身先验输出)被广泛报道为任务能力的证据。我们利用从桌游《Deception: Murder in Hong Kong》抽象出的强制选择信号任务,证明该证据是必要但非充分的。在该环境中,用于判断坐标的参考点(拟合最大化策略、后验最大化策略及均匀随机选择)均可闭式计算。我们在7个语言模型、2个模型族、1次训练后消融实验及3种独立评分规则下开展研究,21个模型-规则组合单元均表现出可靠的项目敏感性。然而,其中8个单元与忽略项目、随机选择的决策器在统计上无显著差异,5个单元在描述目标时的表现差于随机水平。项目敏感性与随机水平的相关系数仅为r=0.30。我们将此现象称为“无对齐的一致性”,并认为它可推广至任何依赖项目敏感性、置换一致性或自一致性且无测量量独立参考的评估场景。此外,我们发现无语用性的字面相似度基线优于多数受测语言模型;在两个基线相似度源上添加语用层会使决策器更趋近随机选择,而非贝叶斯参考;标准标注多项选择格式在此处无可测量的内容信号。所有结果均为预注册工具的模型侧数据,匹配的人类条件已设计并试点,但尚未收集数据。
英文摘要
Item-sensitivity, defined as whether a model's choice depends on the specific input rather than on its own output prior, is widely reported as evidence of task competence. We show this evidence is necessary but not sufficient using a forced-choice signalling task abstracted from the board game Deception: Murder in Hong Kong. In this environment, the reference points against which a coordinate should be judged (a fit-maximising strategy, a posterior-maximising strategy, and uniform random selection) are all computable in closed form. Across seven language models, two model families, a post-training ablation, and three independent scoring rules, every one of 21 model-by-rule cells is reliably item-sensitive. Yet 8 of those 21 cells are not statistically distinguishable from a chooser that ignores the item and selects at random, and 5 score worse than random at describing the target. Item-sensitivity and distance from random correlate at only r = 0.30. We call this consistency without alignment and argue it generalises to any evaluation that relies on item-sensitivity, permutation consistency, or self-consistency without an independent reference for the measured quantity. We further find that a literal-similarity baseline with no pragmatics outperforms most tested language models, that adding a pragmatic layer over two baseline similarity sources moves choosers toward random rather than toward the Bayesian reference, and that a standard labelled multiple-choice format carries no measurable content signal here. All results represent the model side of a pre-registered instrument; a matched human condition is designed and piloted but not yet collected.
Comments13 pages, 3 figures