arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相似模型学习方式不同:最终窗口预训练对训练后行为的塑造超越监督微调

Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT

Cen Lu, Yung-Chen Tang, Andrea Cavallaro

arXiv 2607.25063首次发表:更新:

AI 中文总结

研究探讨模型预训练最后窗口对训练后行为的影响,通过控制实验发现不同预训练最后窗口数据的分支,SFT后表现相近,但后续训练走向不同,安全文本分支有保护效果,表明不能仅依SFT后行为评估模型,还应考虑预训练最后内容。

AI 中文摘要

开发者通过模型行为判断检查点。监督微调(SFT)后,在相关基准测试中表现相近的检查点被视为可互换的。本文探讨这种判断是否忽略了预训练印记,即SFT后基准测试未揭示但决定检查点对进一步训练响应方式的差异。通过在预训练的最后窗口进行控制实验,六个分支从一个部分预训练的检查点分叉,仅在最后窗口的数据不同。SFT后分支表现相近,但相同的训练后阶段,它们在直接偏好优化更新和可验证奖励的强化学习更新下走向非常不同的终点。通过拒绝有害请求来衡量这种偏差,安全文本分支在训练后获得了保护,而其他四个分支几乎没有。保护效果取决于窗口内容,且安全文本需在预训练后期出现。模型预训练的最后内容塑造其对对齐的反应,因此不能仅根据SFT后的行为评估检查点,还应报告其最后训练的数据。

英文摘要

Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant benchmarks are treated as interchangeable, equally ready for the next alignment stage, typically preference optimization. We ask whether this judgment misses a pretraining imprint: a difference that no post-SFT benchmark reveals, yet that decides how each checkpoint responds to further training. To find out, we run a controlled experiment on the final window of pretraining, the last data trained on before instruction tuning. Six branches fork from one partially pretrained checkpoint and differ only in this window: 500 million tokens, 0.1% to 1% of the tokens that precede it. Each branch trains its window on a single data source: generic web text, filtered web text, normative discourse, safety text, mathematical text, or synthetic educational text. SFT and post-training are then identical. After SFT the branches behave near-identically, within about one point on instruction following, refusal, and capability, yet the same post-training carries them to very different endpoints, under both a direct preference optimization update and a reinforcement learning update with a verifiable reward. We measure this deviation through refusal of harmful requests: when post-training begins the safety text branch refuses no more than the web text branch, yet by the end it has lost far less of its refusal. The other four branches gain little or no protection, so the effect is selective to what the window contained. The protection requires the safety text to arrive last rather than earlier in pretraining, and it reproduces on a second model family. What a model is pretrained on last shapes how it reacts to alignment. Therefore, a checkpoint should not be evaluated by its post-SFT behavior alone, and what it was trained on last should be reported with it.

Comments16 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑