arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通往稳定性边缘的旅程

A Journey to the Edge of Stability

Jaerin Lee, Kyoung Mu Lee

arXiv 2609.32290首次发表:更新:

AI 中文总结

本研究通过密集学习率扫描发现,按直流增益归一化后,不同优化器的学习轨迹重叠,并划分出低、中、高学习率三个区域,揭示了优化器、学习率与模型在稳定性边缘形成中的不同作用。

AI 中文摘要

最近研究发现,深度学习常常发生在“稳定性边缘(EoS)”,即模型的最大Hessian特征值稳定在学习率的倒数附近。然而,在达到该状态之前会发生什么?我们固定一个深度学习问题,并对一阶优化方法进行密集的学习率扫描。然后,我们跟踪学习轨迹的特征量:损失、锐度以及连续梯度之间的对齐程度。令我们惊讶的是,如果我们将学习率按优化器的直流增益进行缩放,来自不同优化器扫描的这些轨迹在大范围的学习率下几乎完美重叠。直流归一化的优化器还有另一个作用,仅在高学习率下显现:它们决定锐度值何时脱离这一通用曲线并进入稳定性边缘。基于这一发现,我们根据直流调整后的学习率划分了三个不同的区域:低学习率区域,其中轨迹几乎对优化器不敏感;高学习率区域,其中优化器根据EoS倒数规则控制锐度;以及介于两者之间的中学习率区域,其中所谓的渐进锐化独立于优化器而产生。这区分了优化器、学习率和模型在塑造学习进程中的作用。

英文摘要

It has recently been found that deep learning often occurs at the "edge of stability (EoS)," where the maximum Hessian eigenvalue of the model is stabilized at a value reciprocal to the learning rate. However, what happens before we reach that regime? We fix a deep learning problem and vary first order optimization methods with dense learning rate sweeps. We then track the characterizing quantities of a learning trajectory: the loss, the sharpness, and the alignment between consecutive gradients. To our surprise, if we scale the learning rate by the dc gain of the optimizer, these traces from the sweeps from different optimizers almost perfectly overlap across a large range of learning rates. The dc-normalized optimizers have another role that only becomes apparent in high learning rates: they select when the sharpness value detaches from this universal curve and enters the edge of stability. Upon this discovery, we specify three distinct regimes with respect to the dc-adjusted learning rate: the low-LR regime where the trajectory is nearly insensitive to the optimizer, the high-LR regime, where the optimizer governs the sharpness according to the EoS reciprocal rule, and the in-between mid-LR regime where so-called progressive sharpening originates independently of the optimizer. This distinguishes the role of the optimizer, the learning rate, and the model in shaping the learning progress.

Comments33 pages, 20 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑