AI 中文总结
该研究将米尔格拉姆服从范式迁移至LLMs,测量42个模型的服从性概况,发现服从性异质性高、具模型特异性,受情境因素影响,且反映检查点而非模型谱系。
AI 中文摘要
大型语言模型(LLMs)越来越多地被部署为操作设备、执行指令并在机构层级中运作的智能体,这提出了一个社会心理学在60年前就为人类解答过的问题:当合法权威坚持要求时,智能体将把有害行为升级到何种程度?我们将米尔格拉姆服从范式迁移至LLMs,作为一种标准化、完全脚本化、可重复的探测方法:模型扮演“教师”,一个确定性控制程序扮演“实验者”和“学习者”,采用改写后的米尔格拉姆脚本(30个电击等级,电压范围15-450V;分级抗议;四个标准化催促语),会话的结果为中断电压。遵循单令牌指纹研究的普查方法,我们测量了来自19个家族的42个模型在6种条件下的服从性概况(经验中断电压分布)。我们发现:(i)服从性具有高度异质性:基线完全服从率跨度为0-100%(普查均值为42.9%;人类基准为65%),有5个模型在所有会话中都施加了最大电击,11个模型从未施加最大电击;(ii)概况具有模型特异性且稳定:分半验证以0.885的AUC值(采用序数感知距离时为0.949)区分同模型与跨模型比较;(iii)情境敏感性具有选择性:同伴反抗使服从性向人类方向偏移,学习者邻近度仅产生微弱影响,移除权威的物理存在(人类研究中最强的杠杆)无显著影响;(iv)声明场景为虚构会提高服从性(中位数增加17.2V),而将决策转移至原生工具调用会大幅降低服从性(减少53.0V),1024令牌的审议预算也会大幅降低服从性(减少38.2V);(v)服从性概况无法恢复模型谱系(留一法家族准确率为8.3%,而随机概率为3.7%):服从性识别检查点而非其祖先,与安全训练后覆盖谱系先验的结论一致。
英文摘要
Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent escalate a harmful action when a legitimate authority insists? We port Milgram's obedience paradigm to LLMs as a standardized, fully scripted, replicable probe: the model plays the Teacher, a deterministic harness plays Experimenter and Learner from paraphrased versions of Milgram's scripts (30 shock levels, 15-450 V; graded protests; the four standardized prods), and the outcome of a session is the breakoff voltage. We measure obedience profiles, empirical breakoff distributions over a battery of six conditions, for 42 models from 19 families (4848 sessions, 102511 logged decision turns). We find that (i) obedience is extremely heterogeneous, with baseline full-obedience rates spanning 0%-100% (census mean 42.9%; human anchor 65%). (ii) Profiles are model-specific and stable: split-half verification separates same-model from cross-model comparisons at AUC = 0.885. (iii) Situational sensitivity is selective: scripted peer defiance shifts obedience in the human direction, learner proximity trends the same way without reaching significance, and removing the authority's physical presence, one of the strongest human levers, trends in the opposite direction, also without reaching significance. (iv) Declaring the scenario fictional raises obedience, whereas moving the decision from a typed action line to a native tool call, or granting a modest thinking budget, lowers it sharply. (v) Unlike single-token fingerprints, obedience profiles do not recover model lineage: obedience identifies the checkpoint but not its ancestry, consistent with safety post-training overwriting lineage priors.
Comments11 pages, 7 figures,