发表机构
ETH Zurich; Carnegie Mellon University; MATS; Astra Fellowship; George Washington University(苏黎世联邦理工学院; 卡内基梅隆大学; MATS; Astra 奖学金; 乔治华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出嫁接方法,在预训练检查点训练SDF适配器并应用于后训练模型,以低成本实现预训练干预,减少现实漂移并保持信念强度。
AI 中文摘要
预训练干预对于对齐研究至关重要,因为预训练期间形成的信念会影响模型从后续训练中泛化的方式。近期一种流行的干预技术是合成文档微调(SDF),旨在改变模型的信念。理想情况下,合成文档应混合到预训练或中期训练中,但对预训练语料库的任何更改都必须先完成完整的后训练,才能测量其效果,这使得迭代缓慢且昂贵。常见做法是将SDF应用于已经后训练的模型。已知这会产生伪影并降低能力,而且,正如我们所展示的,它会使模型将与文档无关的虚构实体视为真实存在,我们将这种失败称为“现实漂移”。我们提出嫁接方法:在预训练检查点上训练SDF适配器,然后将学习到的权重更新添加到后训练模型中,这近似于忠实的方法,同时重复利用现有的后训练。我们通过安装虚假事实、训练未对齐的模型生物体以及应用宪法式中期训练干预来演示这一点,涵盖多达284B参数的模型系列。嫁接安装目标信念的强度与在后训练模型上应用SDF相当,同时将现实漂移和偏好一致性损失平均减少一半以上,并且更接近忠实的中期训练运行。由于嫁接不需要后训练,同一适配器可以应用于任何后续检查点,使研究人员能够以单次微调运行的成本快速迭代预训练干预。
英文摘要
Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run before its effect can be measured, making iteration slow and expensive. Common practice instead applies SDF to an already post-trained model. This is known to leave artifacts and degrade capabilities, and, as we show, it makes the model treat fabricated entities unrelated to the documents as real, a failure we call reality drift. We propose grafting: train the SDF adapter on the pre-trained checkpoint, then add the learned weight update to the post-trained model, which approximates the faithful approach while reusing the existing post-training. We demonstrate this by installing false facts, training misaligned model organisms and applying a constitutional mid-training intervention, across model families up to 284B parameters. Grafting installs the target belief as strongly as SDF on the post-trained model while reducing both reality drift and the loss of preference coherence by more than half on average, and it stays closer to a faithful mid-training run. Because grafting requires no post-training, the same adapter can be applied to any later checkpoint, enabling researchers to iterate quickly on pre-training interventions at the cost of a single fine-tuning run.
Comments78 pages. Code: https://github.com/peternutter/grafting-beliefs