非均匀平滑性下Adam的收敛性:与SGDM的可分离性及超越
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
- University of Science and Technology of China(中国科学技术大学)
- Microsoft Research(微软研究院)
- Peking University(北京大学)
- Chinese Academy of Science(中国科学院)
- CUHK (Shenzhen)(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文明确了非均匀平滑性下Adam与SGDM在收敛率上的差异,证明Adam在确定性和随机环境中均能达到一阶优化器的下界,而SGDM存在无法收敛的情况,并引入停时技术进一步优化了收敛率分析。
AI中文摘要:
本文旨在明确区分带动量的随机梯度下降(SGDM)和Adam在收敛率方面的差异。我们证明了在非均匀有界平滑性条件下,Adam比SGDM实现更快的收敛。我们的发现揭示:(1)在确定性环境中,Adam能够达到确定性一阶优化器收敛率的已知下界,而带动量梯度下降(GDM)的收敛率对初始函数值具有更高阶的依赖性;(2)在随机环境中,考虑到初始函数值和最终误差,Adam的收敛率上界与随机一阶优化器的下界相匹配,而存在SGDM在任何学习率下都无法收敛的情况。这些见解明确区分了Adam和SGDM在收敛率上的不同。此外,通过引入一种基于停时的新技术,我们进一步证明,如果考虑迭代过程中的最小梯度范数,相应的收敛率可以匹配所有问题超参数的下界。该技术还可用于证明具有特定超参数调度器的Adam是与参数无关的,因此具有独立的研究价值。
英文摘要:
This paper aims to clearly distinguish between Stochastic Gradient Descent with Momentum (SGDM) and Adam in terms of their convergence rates. We demonstrate that Adam achieves a faster convergence compared to SGDM under the condition of non-uniformly bounded smoothness. Our findings reveal that: (1) in deterministic environments, Adam can attain the known lower bound for the convergence rate of deterministic first-order optimizers, whereas the convergence rate of Gradient Descent with Momentum (GDM) has higher order dependence on the initial function value; (2) in stochastic setting, Adam's convergence rate upper bound matches the lower bounds of stochastic first-order optimizers, considering both the initial function value and the final error, whereas there are instances where SGDM fails to converge with any learning rate. These insights distinctly differentiate Adam and SGDM regarding their convergence rates. Additionally, by introducing a novel stopping-time based technique, we further prove that if we consider the minimum gradient norm during iterations, the corresponding convergence rate can match the lower bounds across all problem hyperparameters. The technique can also help proving that Adam with a specific hyperparameter scheduler is parameter-agnostic, which hence can be of independent interest.