发表机构
EPFL; University of Toronto; Northeastern University; SJTU; Saarland University; Hereon; TUHH; Ontocord AI; TUC; DFKI(洛桑联邦理工学院; 多伦多大学; 东北大学; 上海交通大学; 萨尔大学; 亥姆霍兹极地与海洋研究中心; 汉堡工业大学; Ontocord人工智能公司; 德累斯顿工业大学; 德国人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出合成角色预训练(SPP),在预训练token零阶段植入助手角色,通过标注反思、预训练及角色绑定,提升模型价值构成遵循度与鲁棒性,降低对齐错误率,证明预训练时角色干预是对齐的有效方法。
AI 中文摘要
随着基于语言模型的AI越来越多地部署在自主场景中,使其目标和价值观与人类保持一致变得至关重要。如今,对齐以及助手身份本身通常仅在预训练后引入,此时行为先验已建立,这可能使价值观成为一层薄薄的覆盖,而非根深蒂固,并导致后续的对齐问题。我们追求一种不同的范式,引入合成角色预训练(Synthetic Persona Pretraining,SPP),它在预训练的token零阶段就植入所需的助手角色。首先,我们根据规范价值构成,用与价值对齐的第一人称反思标注预训练文档;其次,我们在标准预训练文档及其反思上通过标准交叉熵损失进行预训练,在众多其他角色中植入所需角色;最后,我们在用户-助手对话数据上进行后训练,将该所需角色与助手身份绑定,此过程我们称为角色绑定。通过在5000亿token上预训练最多30亿参数的模型,我们表明SPP可提升价值构成遵循度和越狱鲁棒性,降低分布外道德困境中的对齐错误率,同时保留模型能力。早期干预很重要:与仅在预训练结束时引入SPP相比,从零开始的对齐会导致价值构成 adherence 更弱,不会改变价值优先级,且在困境中产生的对齐选择更少。这种优势依赖于角色绑定,且重要的是,会随预训练预算增加而提升。总体而言,我们的结果表明早期塑造价值观对齐至关重要,并确立了预训练时的角色干预是实现这一目标的有效方法。
英文摘要
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established. This can make values a thin overlay, rather than deeply rooted, and facilitate subsequent misalignment. Pursuing a different paradigm, we introduce Synthetic Persona Pretraining (SPP), which installs the desired assistant persona from token zero in pretraining. First, we annotate pretraining documents with value-aligned first-person reflections derived from a normative value constitution. Second, we pretrain via the standard cross-entropy loss on standard pretraining documents as well as their reflections, which installs the desired persona among a multitude of other personas. Finally, we post-train on user-assistant dialogue data, which binds this desired persona to the assistant identity, a process we call persona binding. By pretraining models up to 3B parameters on 500B tokens, we show that SPP improves constitution following and jailbreak robustness, and reduces the misalignment rate in out-of-distribution moral dilemmas, while preserving capabilities. Early intervention matters: compared with alignment from token zero, introducing SPP only at the end of pretraining yields weaker constitution adherence, does not shift value priorities, and leads to less aligned choices in dilemmas. This advantage depends on persona binding and, importantly, increases with pretraining budget. Overall, our results show that shaping values early is critical for alignment and establish pretraining-time persona interventions as an effective approach to do so.