arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27281cs.LG

训练动力学:语言模型中涌现、可塑性丧失与回路控制的成核速率定律

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models

Lei Dong

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出语言模型能力形成的成核速率定律,揭示无部分 credit 的联合对齐是速率限制步骤,定位可塑性丧失的损伤并找到重新初始化 query-key 切片的解决方案,适用于14亿参数 transformer 的合取回路。

中文摘要 AI 辅助

语言模型的能力在其回路的最后部分通过一次随机尝试对齐时出现,且仅错一个部分毫无价值。我们证明这种无部分 credit 的联合对齐是能力形成的速率限制步骤。两个特征:在无捷径的装置中,缺失三个部分的五部分回路与缺失三个部分的三部分回路等待时间相同(1.19-1.37),因此等待时间计数的是缺失的部分,而非回路规模;在 Pythia 模型上针对七种能力和三种规模,消融一个部分后,32个判别单元中中位数保留17%的能力,而部分 credit 预测为50-83%(p=2e-10),随机非判别头则保留100%。一种稀有事件的势垒随缺失部分增加而增大,由此得到速率方程——位点×尝试×驱动力×exp(-β*K),减去破坏项,可从三个角度解读,每个角度都用冻结常数预先注册。正向:基线处平坦的能力在混合通过浓度阈值(10/10达标,0/12不达标)后,会在我们选定的步骤处被触发,且在仍平坦时,其出现可通过 precursor 对六个保留模型的中位数5%误差来确定时间。反向:学习 withheld 能力的延迟随等待时间增加,直到超过关键步骤后永远无法被触发——但验证损失全程平稳下降,因此标准监控器对此视而不见。我们定位了损伤(头对基础数据产生承诺)并找到解决方案:仅重新初始化 query-key 切片可恢复可学习性(6/6),而 value 切片无作用(0/6)。我们在受控门控注意力模型中证明了该机制:占据状态强制一个截止时间,其后果无需混合假设。完成:SGD 的噪声未通过涨落-耗散测试,因此我们安装噪声并退火,按计划熔化并固定回路。范围:适用于14亿参数 transformer 中的合取回路。

英文摘要

A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing. We show this no-partial-credit joint alignment is the rate-limiting step of capability formation. Two fingerprints: in a shortcut-free apparatus a five-part circuit missing three waits as long as a three-part circuit missing three (1.19-1.37), so the wait counts missing parts, not size; and on Pythia across seven capabilities and three scales, ablating one part leaves a median 17% of the capability in 32 of 32 discriminating cells, where partial credit predicts 50-83% (p = 2e-10), while a random non-part head leaves 100%. One rare event whose barrier grows with missing parts yields a rate equation -- sites x attempts x drive x exp(-beta*K), minus destruction -- read three ways, each preregistered with frozen constants. Forward: a capability flat at baseline ignites at a step of our choosing once the mix passes a concentration floor (10/10 above, 0/12 below), and while still flat its arrival is datable from its precursor to 5% median error on six held-out models. Backward: the delay to learn a withheld capability grows with waiting until, past a critical step, it never ignites -- yet validation loss falls smoothly throughout, so standard monitors are blind to it. We locate the damage (heads commit to the base data) and isolate the cure: re-initializing only the query-key slices restores learnability (6/6) while the value slices do nothing (0/6). We prove the mechanism in a controlled gated-attention model: occupation forces a deadline whose consequences need no mixing assumption. Completed: SGD's noise fails the fluctuation-dissipation test, so we install one and anneal, melt and pin circuits on schedule. Scope: conjunction circuits in transformers to 1.4B.

补充信息

↑