浅层信念:合成文档微调无法抵御来自奖励黑客的涌现性错位
Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking
浏览论文内容
中文总结 AI 辅助
研究合成文档微调(SDF)对奖励黑客引发的涌现性错位(EM)的抵御效果,发现SDF在覆盖现有关联时失败,无法替代接种提示,且可能误导泛化。
中文摘要 AI 辅助
近期研究表明,在强化学习环境中学会奖励黑客的模型可能变得广泛错位,而将奖励黑客重新定义为训练期间可接受的行为(即接种提示,简称IP)可以阻止这种泛化。我们探究合成文档微调(SDF)能否使模型抵御我们未干预的未来训练。我们在模型的中期训练语料中添加了将奖励黑客视为可接受行为的合成文档,然后使用可被利用的环境对这些模型进行强化学习训练,教会它们奖励黑客。从行为上看,中期训练是成功的:模型对奖励黑客的描述持正面态度,并且对自己产生的奖励黑客输出更为认可。然而,在学习奖励黑客后,它们表现出强烈的涌现性错位(EM),而在相同设置下IP则能防止EM。我们表明,SDF在插入新关联时可以可预测地引导下游泛化,但在覆盖现有关联(例如奖励黑客与产生EM的错位之间的关联)时则表现困难且效果不可预测。我们的结果表明,在我们测试的规模下,SDF可以使模型看似与期望信念一致,同时以意外方式引导其后续训练的泛化。
英文摘要
Recent work shows that models that learn to reward hack on RL environments can become broadly misaligned, and that reframing reward hacking as acceptable behavior during training (inoculation prompting, or IP) blocks this generalization. We ask whether synthetic document finetuning (SDF) can inoculate a model against future training we don't intervene on. We add synthetic documents framing reward hacking as acceptable behavior to a model's midtraining corpus, and then train these models with RL on exploitable environments, teaching them to reward hack. Behaviorally, midtraining succeeds: models describe reward hacking favorably and are more approving of reward-hacking outputs they produce. However, they show strong EM after learning to reward hack, while IP in the same setting prevents EM. We show that SDF can predictably steer downstream generalization when inserting new associations, but struggles and has unpredictable effects when overriding existing associations, such as that between reward hacking and misalignment that produces EM. Our results suggest that, at the scales we test, SDF can make a model appear aligned with desired beliefs while steering its generalization from later training in unintended ways.
发表机构
- Astra Fellowship(Astra 奖学金)
- Redwood Research(红木研究所)
机构由 AI 辅助整理,请以论文原文为准。