Adam 降低了一种独特的锐度形式:极小值流形附近的理论洞察
Adam Reduces a Unique Form of Sharpness: Theoretical Insights Near the Minimizer Manifold
- Tsinghua University(清华大学)
- Institute for Interdisciplinary Information Sciences(交叉信息研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文证明Adam优化器隐式降低由其自适应更新塑造的独特锐度度量,在极小值流形附近表现出与SGD性质不同的优化行为,并在标签噪声和稀疏回归场景中展现出更好的稀疏性与泛化能力,为自适应梯度方法提供了统一分析视角。
AI中文摘要:
尽管 Adam 优化器在实践中广受欢迎,但大多数理论分析仍将随机梯度下降(SGD)作为 Adam 的替代研究对象,因此人们对 Adam 所找到的解与 SGD 有何不同知之甚少。在本文中,我们证明 Adam 隐式地降低了一种由其自适应更新所塑造的独特锐度度量,从而得到与 SGD 在性质上不同的解。更具体地说,当训练损失很小时,Adam 会在极小值流形附近游走,并以半梯度方式自适应地最小化该锐度度量,我们通过使用随机微分方程的连续时间近似严格刻画了这一行为。我们进一步在一个已被充分研究的场景中展示了这种行为与 SGD 的差异:当使用标签噪声训练过参数化模型时,已有研究表明 SGD 会最小化 Hessian 矩阵的迹 $\tr(\mH)$,而我们证明 Adam 转而最小化 $\tr(\Diag(\mH)^{1/2})$。在求解对角线性网络的稀疏线性回归问题时,这一区别使 Adam 能够获得比 SGD 更好的稀疏性和泛化性能。最后,我们的分析框架不仅适用于 Adam,还扩展到一类广泛的自适应梯度方法,包括 RMSProp、Adam-mini、Adalayer 和 Shampoo,并为这些自适应优化器如何降低锐度提供了统一的视角,我们希望这能为未来的优化器设计提供启示。
英文摘要:
Despite the popularity of the Adam optimizer in practice, most theoretical analyses study Stochastic Gradient Descent (SGD) as a proxy for Adam, and little is known about how the solutions found by Adam differ. In this paper, we show that Adam implicitly reduces a unique form of sharpness measure shaped by its adaptive updates, leading to qualitatively different solutions from SGD. More specifically, when the training loss is small, Adam wanders around the manifold of minimizers and takes semi-gradients to minimize this sharpness measure in an adaptive manner, a behavior we rigorously characterize through a continuous-time approximation using stochastic differential equations. We further demonstrate how this behavior differs from that of SGD in a well-studied setting: when training overparameterized models with label noise, SGD has been shown to minimize the trace of the Hessian matrix, $\tr(\mH)$, whereas we prove that Adam minimizes $\tr(\Diag(\mH)^{1/2})$ instead. In solving sparse linear regression with diagonal linear networks, this distinction enables Adam to achieve better sparsity and generalization than SGD. Finally, our analysis framework extends beyond Adam to a broad class of adaptive gradient methods, including RMSProp, Adam-mini, Adalayer and Shampoo, and provides a unified perspective on how these adaptive optimizers reduce sharpness, which we hope will offer insights for future optimizer design.