发表机构
Aimpoint Digital Labs; Mistral AI; New York University(Aimpoint Digital实验室; Mistral AI公司; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过训练轨迹分析与受控干预,揭示大规模激活(注意力汇聚点)的生命周期,发现权重衰减因果控制其规模更替,并建立平衡模型,确立其为训练时调节激活幅度的关键杠杆。
AI 中文摘要
大规模激活(即残差流中幅度远大于典型激活的坐标)与Transformer中的注意力汇聚点相关,但其规模在训练过程中如何被调节仍未完全理解。结合训练轨迹分析与受控干预,我们追踪了它们的出现、增长与巩固过程。携带汇聚点的通道在不同随机种子间有所变化,但在每次运行内部早期趋于稳定。随着训练的延长,周围通道被侵蚀,汇聚点集中于少数冗余的载体上。在消融实验中,梯度衰减跟随汇聚点标记的集体均方根幅度而非任何单一通道,这使得集体规模成为理解其影响的核心。我们的核心结果是,权重衰减因果性地控制全局激活规模的更替。在受控延续训练中,在峰值附近移除衰减会使规模持续上升,而保留衰减即使在恒定学习率下也会导致规模下降。我们开发了一个用于大规模激活幅度上升与峰值的平衡模型,其中AdamW预条件化的增长与权重衰减相对抗。扫描衰减系数$\lambda$使峰值时间近似对数线性移动,并产生近似为$\lambda^{-1/2}$缩放的峰值幅度,这与该平衡一致。优化器测量进一步表明,预条件化在原始维持力过小的情况下,仍能维持大通道群体对抗衰减。综合来看,这些发现将观察到的生命周期与规模调节的训练动态联系起来,并确立权重衰减作为训练时对激活幅度的控制杠杆。
英文摘要
Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-trajectory analyses and controlled interventions, we trace their emergence, growth, and consolidation. Sink-carrying channels vary across random seeds but stabilize early within each run. Over longer training, surrounding channels erode and the sink concentrates onto a few redundant carriers. Across ablations, gradient attenuation follows the sink token's collective root-mean-square magnitude rather than any single channel, making collective scale central to understanding their effects. Our central result is that weight decay causally controls the turnover of global activation scale. In controlled continuations, removing decay near the peak allows this scale to keep rising, whereas retaining it produces decline even at constant learning rate. We develop a balance model for the rise and peak of massive-activation magnitude, in which AdamW-preconditioned growth opposes weight decay. Sweeping the decay coefficient $λ$ shifts peak timing approximately log-linearly and yields peak magnitudes scaling approximately as $λ^{-1/2}$, consistent with this balance. Optimizer measurements further show that preconditioning sustains the large-channel cohort against decay even when raw maintaining forces are too small to do so. Together, these findings connect the observed life cycle to scale-regulating training dynamics and establish weight decay as a training-time lever on activation magnitude.