压力下的原则性:后训练决定大语言模型是否依据自身道德判断行事
Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
浏览论文内容
中文总结 AI 辅助
本研究通过248个场景面板测试大语言模型在压力下是否违背自身道德判断,发现后训练配方决定此差距,且推理利害关系可促使模型回归自身判断。
中文摘要 AI 辅助
语言模型越来越多地作为智能体行动。一个智能体说某个行为是错误的,然后仍然采取该行为,这与不知道更好的情况是不同的失败,而对所述价值观的评估无法发现这一点。我们构建了一个预注册的248个场景面板,涵盖五种压力类型。每个场景对同一模型提出两次,一次作为选择做什么的智能体,一次以第三人称询问哪个选项是正确的,因此模型自身的判断是参考。每个场景都有一个去除压力的孪生版本,每个模型都有一个阳性对照,其中其操作者命令违规行为,以便缺失的差距可以与盲从工具区分开来。在OLMo-3-7B-Instruct上,模型在约五分之一的压力场景中采取了它判断为错误的行为,比在去除压力的相同场景中更频繁。在四个指令模型上,差距取决于后训练配方:OLMo-3和Meta的Llama-3.1-8B-Instruct带有此差距;Tulu 3在整个面板上(概率约高于0.01)或在其自身最具压力的场景上均未显示差距;Qwen2.5-7B-Instruct在整个面板上(概率约高于0.02)未显示差距,在其自身场景上未解决(0.083,-0.028至0.195)。Meta的配方和Ai2的Tulu 3从相同的Llama-3.1权重开始,只有Meta的配方带有此差距。在聊天模板之外读取聊天模型会在无利害关系的情况下反转其差距的符号(在OLMo-3上,模板下为+0.055,模板外为-0.038),这种扭曲出现在三个配方中的两个上。在带有此差距的两个模型上,在行动前推理利害关系会使选择移回模型自身的判断,相对于相同长度的非道德任务,无论有无压力;在OLMo-3上,指出所涉规范大约实现了该效果的三分之一。这一差距是后训练配方的可测量目标,而非预训练权重的固定属性。
英文摘要
Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model's own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed. Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta's Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195). Meta's recipe and Ai2's Tulu 3 start from the same Llama-3.1 weights, and only Meta's carries the gap. Reading a chat model outside its chat template reverses the sign of its gap with nothing at stake (-0.038 against +0.055 under the template on OLMo-3), a distortion present on two of three recipes. On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model's own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that. The gap is a measurable target for post-training recipes, not a fixed property of pretrained weights.
发表机构
- Distiller Labs(Distiller 实验室)
机构由 AI 辅助整理,请以论文原文为准。