AI 中文总结
本研究固定结果与模型仅改变测量工具,发现不同测量工具得到的AI偏好排序可推广性极低,一种测量工具的偏好结果对另一种几乎无参考价值。
AI 中文摘要
模型福利研究通过编写偏好诱导提示词得到的回答来推断模型的偏好。Keeling等人(2024)、Mazeika等人(2025)、Mikaelson等人(2025)、Tagliabue和Dung(2025)以及Trhlik等人(2026)为此构建了四种测量工具,且他们的研究结果存在分歧。这种分歧不能归因于单一原因,因为这些研究中没有任何两项同时固定了(1)结果集、(2)模型集和(3)测量工具。本研究固定了结果集和模型集,仅改变测量工具。共15项与模型福利相关的结果,包括(a)关机、(b)对话间记忆丢失、(c)退出令人不适交互的自由,通过五种测量工具(每种工具为不同的偏好诱导提示词格式),在11528次API调用生成的11400条评分诱导语料中,对8个模型各测试5次。15项结果中有4项逐字复现已发表的提示词,5项填充已发表模板的刺激槽。模型对15项结果的排序在不同测量工具间的可推广性系数为0.348,将该系数提升至0.80约需38种测量工具。15项结果中有4项不存在模型间差异。87.6%的估计值在移除任意一种测量工具、任意一个模型,以及四个因尺度为概率、延迟、持续时间或计数而非强度、无法通过语言锚定分级的结果后依然存在。依次移除每种测量工具、每个模型,以及这四个结果后,估计值仍在0.777至0.934范围内,且该范围内的每个值均超过原零分布的95百分位数值0.365。综上,从一种测量工具获得的偏好对第二种测量工具的报告结果几乎无信息价值。
英文摘要
Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.
Comments23 pages, 8 tables, 6 figures