发表机构
Mohamed bin Zayed University of Artificial Intelligence; Aerospace Information Technology University; The Ohio State University; JD.com, Inc.; Centre for Artificial Intelligence and Robotics; Hong Kong Institute of Science & Innovation; Chinese Academy of Sciences(穆罕默德·本·扎耶德人工智能大学; 航天信息技术大学; 俄亥俄州立大学; 京东集团; 人工智能与机器人中心; 香港科技创新研究院; 中国科学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明在广义光滑性下,仅需随机梯度的第二矩信息即可保证Adam算法收敛,无需强尾部假设,并给出了高概率与期望收敛速率及最优置信度依赖。
AI 中文摘要
Adam算法被广泛观察到即使在目标函数显著偏离全局光滑性时也能保持稳定。然而,在广义光滑性框架下,现有的分析依赖于对随机梯度的强尾部假设,例如几乎必然有界或次高斯性。Li等人(2023)指出,在仅利用随机梯度的第二矩信息且无此类集中性假设的情况下,Adam是否能在广义光滑目标上收敛是一个重要的开放方向。本文在相当一般的条件下给出了肯定回答:此类尾部假设并非必要。基于Jin等人(2026)为经典光滑性和有界方差开发的Adam自归一化框架,我们将停时和去预处理策略扩展到$L_0$-$L_p$广义光滑性条件和广义第二矩ABC条件。即使随机梯度条件仅提供可能沿轨迹增长的第二矩信息,Adam的随机轨迹仍保持在局部良态的光滑性区域,并在有界方差和全局光滑性下具有拉伸指数尾部衰减。因此,我们在$p<2$的完整范围内建立了高概率收敛速率保证,置信度依赖为$\u03b4^{-1/2}$阶,而步长前因子对$\u03b4$的依赖仅通过单个对数因子。我们进一步构造了一个困难实例,表明在仅第二矩信息下,这种$\u03b4^{-1/2}$型置信度依赖是尖锐的。最后,在$p<1$的范围内,我们将轨迹控制与稀有事件的多项式增长估计相结合,获得了期望意义上的收敛速率保证。
英文摘要
Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whether Adam converges on generalized smooth objectives under only second moment information on the stochastic gradients, without such concentration assumptions, was identified as an important open direction by Li et al. (2023). This paper gives an affirmative answer under fairly general conditions: such tail assumptions are not necessary. Building on the Adam self-normalization framework of Jin et al. (2026), developed for classical smoothness and bounded variance, we extend the stopping-time and de-preconditioning strategy to the $L_0$-$L_p$ generalized smoothness condition and a generalized second moment ABC condition. Even when the stochastic-gradient condition provides only second moment information that may grow along the trajectory, the stochastic trajectory of Adam remains in a locally well-behaved smoothness region, with stretched-exponential tail decay under bounded variance and global smoothness. Consequently, we establish high-probability convergence rate guarantees over the full range $p<2$, with confidence dependence of order $δ^{-1/2}$, while the stepsize prefactor depends on $δ$ only through a single logarithmic factor. We further construct a hard instance showing that, under only second-moment information, this $δ^{-1/2}$-type confidence dependence is sharp. Finally, in the regime $p<1$, we combine the trajectory control with polynomial-growth estimates on rare events to obtain convergence rate guarantees in expectation.
Comments37 pages, 4 figures. Accepted at NeurIPS 2026