arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20412cs.CLcs.AI

压力测试对齐中训练

Stress-testing Alignment Midtraining

  • Arcadia Impact(阿卡迪亚影响力)
  • Resolution

机构由 AI 辅助整理,请以论文原文为准。

Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan

AI总结:

本研究通过大规模实验(最高1100亿参数模型、10亿中训练token)压力测试对齐中训练(AMT),发现其效果脆弱:少量竞争性微调数据即可抹除其作用,且规则演示必须存在于中训练或后训练数据中,故现有证据不足以支持AMT能解决对齐核心难题。

AI中文摘要:

在通过后训练技术对齐前沿模型时,不可能直接展示我们希望模型在所有可能的部署环境中表现出的所有行为;我们的模型必须在后训练分布之外进行泛化。一种提出的解决方案是对齐中训练(AMT),它继续在大量与对齐相关的文档上进行预训练,以鼓励在训练的后期阶段实现泛化。尽管AMT作为一种对齐方法备受关注,但关于其有效性的公开证据有限。为解决这一问题,我们识别了围绕中训练的若干假设,并在不同规模下对其进行评估:模型规模高达1100亿参数,中训练数据量达10亿个token。例如,我们研究了一种场景,其中后训练数据在两种可能的动机之间是模糊的。我们发现,在这种设置的简单版本中,中训练可以引导模型的动机。然而,存在一小部分微调数据暗示了竞争性动机,这抹去了AMT的效果。我们还研究了我们希望AI遵循若干规则但仅演示其中一部分的场景。我们发现,要使这些规则被稳健地学习,演示必须存在于中训练或后训练数据集中。基于这些及其他发现,我们认为目前没有足够的公开证据让我们自信地断言,中训练能够解决对齐强大AI系统所固有的核心困难。

英文摘要:

When aligning frontier models through post-training techniques, it is not possible to directly demonstrate all of the behaviours we want a model to exhibit in all possible deployment environments; our model must generalise outside of the post-training distribution. One proposed solution is alignment midtraining (AMT), which continues pretraining on large volumes of alignment-relevant documents to encourage generalisation in later stages of training. Despite the prominence of AMT as an alignment approach, there is limited public evidence for its effectiveness. To resolve this, we identify several assumptions around midtraining and evaluate them across scale: up to 110 billion-parameter models and 1 billion midtraining tokens. For instance, we study a scenario where post-training data is ambiguous between two possible motivations. We find that midtraining can steer the model's motivation in simple versions of this setting. However, the presence of a tiny fraction of finetuning data which suggests a competing motivation erases the effects of AMT. We also study scenarios in which we want an AI to follow a number of rules, but only demonstrate a subset of them. We find that demonstrations must be present either in midtraining or post-training datasets for these rules to be robustly learned. Based on these and other findings, we do not believe that there is sufficient public evidence for us to confidently state that midtraining can address the core difficulties inherent in aligning powerful AI systems.

↑