arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33150cs.CL

语言模型预训练中的泛化动力学

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建评测套件揭示语言模型预训练中存在“模式跳跃”现象,即模型在泛化与记忆间反复切换,并将其归因于容量分配问题,同时提出利用该现象选择检查点与数据以增强泛化。

中文摘要 AI 辅助

人们通常假设,在预训练过程中,语言模型(LM)会从模式匹配的“鹦鹉”稳定地成熟为具备可泛化智能的实体。我们构建了一个小型评测套件,并证明这种心智模型是错误的:在整个预训练过程中,语言模型会频繁且突然地在“鹦鹉式”计算与“智能式”计算之间跳跃。我们将这种现象称为“模式跳跃”(mode-hopping)。在我们的评测套件中,语言模型会突然固守于记忆化或上下文中的模式,而非进行上下文学习;使用系统1思维而非系统2思维;采纳听起来正确的内容而非真正正确的内容;在多跳人物问答、上下文外推理中失败,并出现突发的错位(emergent misalignment)——随后又同样突然地恢复并表现出泛化能力。模式跳跃无法用标准优化动力学解释:它在局部是稳定的,且无法通过检查点平均(checkpoint averaging)来修复。我们转而将其视为一种容量分配问题:在容量受限的模型中,可泛化的电路必须与训练早期学到的浅层电路竞争,而每个预训练窗口中的数据可能决定哪些电路胜出。我们的评测套件为泛化研究提供了一种新的高效视角。我们展示了两个具体应用:(i)选择在推理和一致性上强泛化的中间预训练检查点,其效果优于最终的预训练或中期训练检查点;(ii)选择能够控制和稳定泛化动力学的预训练数据。

英文摘要

People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like computations. We call this mode-hopping. Across our suite, LMs suddenly latch onto memorized or in-context patterns instead of in-context learning, use System 1 instead of System 2 thinking, pick up what sounds true instead of what is true, fail at multi-hop persona QA, out-of-context reasoning, and emergent misalignment -- then just as suddenly revert and generalize. Mode-hopping is not explained by standard optimization dynamics: it is locally stable and cannot be fixed by checkpoint averaging. We instead think of it as a capacity allocation problem: in a capacity-bounded model, generalizable circuits must compete with the shallow ones learned early in training, and the data in each pre-training window may decide which circuits win. Our suite provides a new efficient lens on generalization. We demonstrate two concrete applications: (i) select intermediate pre-training checkpoints that strongly generalize reasoning and alignment, better than the final pre- or mid-training checkpoints, and (ii) select pre-training data that controls and stabilizes generalization dynamics.

补充信息

↑