AI 中文总结
研究热身 - 稳定 - 衰减学习率调度中冷却阶段效果,发现其取决于梯度噪声结构和优化器更新归一化。通过理论分析和模拟,揭示不同情况下冷却的作用机制,还在真实分类任务中验证噪声 regime 诊断,解释冷却是否有帮助。
AI 中文摘要
热身 - 稳定 - 衰减(WSD)学习率调度的冷却阶段,如今在大模型预训练中是默认设置,在某些情况下能降低最终训练损失,而在其他情况下则毫无作用。我们给出了能证明哪种情况会出现的依据,这取决于两个属性:梯度噪声的结构以及优化器是否对其更新进行归一化。在具有乘法(梯度成比例)噪声的强凸目标上,随机梯度下降以恒定学习率几何收缩,所以冷却并无改善之处。在相同目标和噪声下,基于符号和归一化方法(自适应优化器的标准替代方法)会稳定在\(\eta^2\)阶的噪声下限,并且只有当学习率趋近于零时才会达到最小值;然后任何加性噪声都会为每种方法恢复一个下限。其机制很简单:随机梯度下降步骤与梯度成比例收缩,从而自我退火,而归一化步骤保持单位尺度,无法做到。我们精确求解了二次函数上符号随机梯度下降的平稳定律,并以封闭形式得到下限常数,在\((L_0,L_1)\)平滑性下证明了局部解离形式,通过尺度不变性论证将下限扩展到\(d>1\)维的归一化随机梯度下降,并建立了对动量和重尾噪声的鲁棒性。模拟证实了每一个预测,并且我们在一个直接测量梯度噪声的真实分类任务中展示了由此产生的噪声 regime 诊断。该机制解释了冷却是否有帮助;在尺度上使用的内部冷却分数位于平稳景观和噪声几何之外。
英文摘要
The cooldown phase of a warmup-stable-decay (WSD) learning-rate schedule, now a default in large-model pretraining, lowers the final training loss in some settings and does nothing in others. We give a provable account of which case obtains, and it turns on two properties together: the structure of the gradient noise and whether the optimizer normalizes its update. On a strongly convex objective with multiplicative (gradient-proportional) noise, stochastic gradient descent contracts geometrically at a constant learning rate, so cooldown has nothing to improve. Under the same objective and noise, sign-based and normalized methods, the standard surrogates for adaptive optimizers, settle on a noise floor of order $η^2$ and reach the minimizer only as the learning rate is driven to zero; any additive noise then reinstates a floor for every method. The mechanism is elementary: an SGD step shrinks in proportion to the gradient and so anneals itself, whereas a normalized step keeps unit scale and cannot. We solve the signSGD stationary law on the quadratic exactly and obtain the floor constant in closed form, prove a local form of the dissociation under $(L_0,L_1)$-smoothness, extend the floor to normalized SGD in dimension d>1 by a scale-invariance argument, and establish robustness to momentum and heavy-tailed noise. Simulation confirms every prediction, and we demonstrate the resulting noise-regime diagnostic on a real classification task with directly measured gradient noise. The mechanism explains whether cooldown helps; the interior cooldown fraction used at scale lies outside stationary landscape-and-noise geometry.
Comments11 pages, 12 figures