发表机构
Indian Institute of Science Education and Research Kolkata; The University of Manchester(印度科学教育与研究学院加尔各答分校; 曼彻斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文在Hölder光滑和重尾噪声下,证明标准SGD、δ-GClip和G-Clip的收敛速率,其中G-Clip在极重尾区域首次获得收敛保证。
AI 中文摘要
随机梯度方法的经典收敛性保证通常假设目标函数是Lipschitz光滑的且梯度噪声具有有限方差,而这些假设在实践中经常被违反。相反,我们研究在这些假设联合放宽下的非凸随机优化:目标函数具有$(L,s)$-Hölder连续梯度,其中$s\in(0,1]$,且梯度噪声仅满足有界$\alpha$阶矩条件,其中$\alpha\in(1,2]$。我们建立了三个收敛结果。首先,当$\alpha\ge1+s$时,标准SGD以$O(T^{-s/(1+s)})$的速率收敛,将经典非凸SGD速率同时扩展到重尾噪声和Hölder光滑情形。其次,我们分析了$\delta$正则化梯度裁剪($\delta$-GClip),一种可证明训练宽深网络的训练器,并在相同条件下建立了$O(T^{-2s(\alpha-1)/[(1+s)(2\alpha-1)]})$的平稳性速率。第三,我们分析标准梯度裁剪(G-Clip),并表明当$\alpha\ge1+s$时它恢复上述速率,而在极重尾区域$\alpha<1+s$中,其收敛速率为$O(T^{-2s(\alpha-1)/[(\alpha-1)+s(2\alpha-1)]})$——这是该区域中任何基于随机梯度方法的首个收敛保证。
英文摘要
Classical convergence guarantees for stochastic gradient methods typically assume Lipschitz-smooth objectives and finite-variance gradient noise, both frequently violated in practice. In contrast, we study nonconvex stochastic optimization under the joint relaxation of these assumptions: objectives with $(L,s)$-Hölder continuous gradients, $s\in(0,1]$, and gradient noise satisfying only a bounded $α$-th moment condition for $α\in(1,2]$. We establish three convergence results. Firstly, that standard SGD converges at rate $O(T^{-s/(1+s)})$ whenever $α\ge1+s$, extending the classical nonconvex SGD rate to heavy-tailed noise and Hölder smoothness simultaneously. Secondly, we analyze $δ$-regularized gradient clipping ($δ$-GClip), a provable trainer of wide and deep nets, and establish a stationarity rate of $O(T^{-2s(α-1)/[(1+s)(2α-1)]})$ under the same condition. Thirdly, we analyze standard gradient clipping (G-Clip) and show that it recovers the above rate for $α\ge1+s$ while in the very heavy-tailed regime $α<1+s$, it has a convergence rate $O(T^{-2s(α-1)/[(α-1)+s(2α-1)]})$ --- the first convergence guarantee in this regime for any stochastic gradient based method.
Comments46 pages, 2 figures