arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应学习率重缩放的可解模型:加速、稳定性与缩放

A Solvable Model of Adaptive Learning Rate Rescaling: Acceleration, Stability & Scaling

Itay Lavie, Clarissa Lauditi, Cengiz Pehlevan

arXiv 2610.06701首次发表:更新:

发表机构

Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过可解模型研究归一化SGD中自适应学习率重缩放,揭示其加速与边际稳定性机制,并量化宽度和批量对资源缩放的影响。

AI 中文摘要

现代优化器中的一个常见设计原则是将更新幅度与原始梯度范数解耦,但其对学习曲线和资源缩放的影响仍不明确。我们通过研究随机特征模型中的归一化随机梯度下降(SGD)来隔离这一机制,其中教师和数据协方差具有幂律分布。固定范数更新会诱导有效学习率随梯度缩小而增大。我们推导了一个动力学平均场理论(DMFT),描述了损失对训练时间、模型宽度和批量大小的联合依赖。归一化最初加速SGD,将幂律指数$r_{\ m SGD}<1$映射为$2r_{\ m SGD}/(1-r_{\ m SGD})$,在$r_{\ m SGD}=1$时实现指数收敛,在$r_{\ m SGD}>1$时实现形式上的有限时间收敛。然而,在有限步长下,同样的反馈最终会破坏加速并导致边际稳定性。后期理论产生了宽度受限、随机稳定性边缘(EoSS)和确定性稳定性边缘(EoS)三个区域。这些阶段决定了在可比计算量下,更大的批量或更宽的模型何时能减少串行训练时间。我们量化了在这些阶段中,增加批量大小或宽度能否通过更少的优化步骤来补偿每步的额外计算,以达到目标损失。在CIFAR-5M上的线性化ResNet实验支持了预测的加速、破坏和资源缩放趋势。总之,这些结果在一个可解理论中连接了归一化诱导的加速、EoS效应和宽度-批量分配。

英文摘要

A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, yet its consequences for learning-curve and resource scaling remain unclear. We isolate this mechanism by studying normalized SGD in a random-feature model with power-law teacher and data covariance. Fixed-norm updates induce an effective learning rate that grows as gradients shrink. We derive a dynamical mean-field theory (DMFT) describing the joint dependence of the loss on training time, model width and batch size. Normalization initially accelerates SGD, mapping the power-law exponent $r_{\rm SGD}<1$ to $2r_{\rm SGD}/(1-r_{\rm SGD})$, with exponential convergence at $r_{\rm SGD}=1$ and formal finite-time convergence for $r_{\rm SGD}>1$. At finite step size, however, the same feedback ultimately breaks the acceleration and leads to marginal stability. The late-time theory yields width-limited, edge-of-stochastic-stability (EoSS), and deterministic edge-of-stability (EoS) regimes. These phases determine when larger batches or wider models reduce serial training time at comparable compute. We quantify in which of these phases increased batch size or width can compensate the excess compute use per step by fewer optimization steps to target loss. Linearized ResNet experiments on CIFAR-5M support the predicted acceleration, breakdown, and resource-scaling trends. Together, these results connect normalization-induced acceleration, EoS effects, and width--batch allocation within a solvable theory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑