arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

宪法式中间训练:内容存在驱动对齐增益

Constitutional Midtraining: Content Presence Drives Alignment Gains

Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel Shadbolt

arXiv 2607.26654首次发表:更新:

发表机构

University of Oxford; Institute for Ethics in AI, University of Oxford; Geodesic Research; Oxford Internet Institute, University of Oxford(牛津大学; 牛津大学人工智能伦理研究所; Geodesic Research; 牛津大学互联网研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出宪法式中间训练方法,在中间训练中插入基于原则价值观的内容,可提升模型对齐的泛化性与持久性,削弱SFT灌输的敲诈倾向且不造成能力损失,为以SFT为中心的流程提供低成本补充方案。

AI 中文摘要

后训练对齐往往较为肤浅,在微调过程中会被削弱。与后训练清晰分离的中间训练干预措施能否产生持久的对齐效果,这一问题尚未得到验证。我们通过宪法式中间训练对此展开测试:在中间训练中插入基于原则、价值观的内容,以1200亿参数规模的仅重放控制组作为对照。我们基于Anthropic的宪法构建了3.94亿token的宪法语料库,采用2×2因子设计(课程顺序×审议推理),生成四种宪法式中间训练条件及一个控制组,在自生成及既定基准上进行评估,包括压力下的对齐、价值冲突解决、敲诈以及三个阶段(中间训练后、SFT后、良性微调后)的涌现失准。宪法式中间训练模型在对齐泛化性和持久性上优于控制组,尤其在敲诈任务中:SFT会在所有模型中灌输敲诈倾向,但宪法式中间训练可削弱该倾向,且这一优势在良性微调后仍存在(降低17.5个百分点)。这一优势并未延伸至需要主动抵抗上下文压力或冲突的场景,在SFT后优势会减弱。中间训练中宪法内容的存在比其结构更重要,且在我们测试的所有阶段(MMLU、ARC-Easy、piqa、GSM8K),宪法式中间训练在平均水平上不会造成能力损失。因此,中间训练中适量的宪法内容可产生广泛、持久的对齐增益,为以SFT为中心的流程提供一种低成本的补充方案。代码、数据和模型均已公开。

英文摘要

Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtraining at 120B scale, where principled, values-based content is inserted into midtraining. A 2x2 design (curriculum ordering x deliberative reasoning) was used to produce four constitutionally midtrained conditions, plus a control, which were evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment. All models were evaluated across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperformed the control on alignment generalization and durability, notably on blackmail: SFT instilled a blackmail propensity in all models, but constitutional midtraining blunted it, with the advantage surviving benign fine-tuning (-17.5pp). This durability did not extend to settings that required active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also mattered more than its structure, and constitutional midtraining incurred no capability cost, on average, at any stage (MMLU, ARC-Easy, piqa, GSM8K). A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑