发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究常步长随机梯度下降在平坦极小值处的缩放极限,通过分析由压缩驱动链生成马尔可夫噪声的SGD,证明其在特定距离下的收敛性,给出不同平坦指数下的收缩因子及小步长缩放极限,揭示了与强凸情况不同的行为。
AI 中文摘要
对于具有常步长α的随机梯度下降(SGD),以极小值点为中心的迭代不变律描述了算法在长时间范围内的行为。在强凸情况下,该不变律具有熟悉的√α缩放,且当α↓0时极限为高斯分布。我们表明,对于具有平坦极小值和(次)二次尾部的凸目标H,这种行为会发生根本变化。具体而言,我们研究由压缩驱动链生成马尔可夫噪声的SGD。对于每个足够小的常步长α,我们证明了在由α相关度量诱导的Wasserstein距离中,存在、唯一且几何收敛到一个增强的不变律。当极小值点x*具有局部平坦指数m≥2时,我们得到收缩因子为1 - cα^(m - 1),在二次情况m = 2时恢复为1 - cα。然后我们分析小步长缩放极限。我们表明不变律集中在α^(1/m)尺度上,并且重新缩放后的迭代弱收敛到随机微分方程dY_t = -h_0(Y_t)dt + Σ^(1/2)dB_t的平稳分布,其中h_0是极小值点处的极限漂移,Σ表示渐近协方差。当m = 2时恢复高斯极限,在平坦情况m>2时通常给出非高斯平稳极限。最后,我们给出了具有不等平坦指数的坐标可分目标的相应结果。
英文摘要
For stochastic gradient descent (SGD) with a constant stepsize $α$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons. In the strongly convex case, this invariant law has the familiar $\sqrtα$ scaling and a Gaussian limit as $α\downarrow 0$. We show that this behavior changes fundamentally for convex objectives $H$ with flat minima and (sub)quadratic tails. More specifically, we study SGD with Markovian noise generated by a contractive driving chain. For every sufficiently small constant stepsize $α$, we prove existence, uniqueness, and geometric convergence to an augmented invariant law in a Wasserstein distance induced by an $α$-dependent metric. When the minimizer $x_\star$ has local flatness exponent $m\ge2$, meaning that $\nabla^2 H(x)\asymp \lVert x-x_\star\rVert^{m-2} I_d$ as $x\to x_\star$, we obtain a contraction bound with factor $1-cα^{m-1}$, where $c>0$ is a constant. This recovers the factor $1-cα$ in the quadratic case $m=2$. We then analyze the small-stepsize scaling limit. We show that the invariant law concentrates on the scale $α^{1/m}$ and that the rescaled iterates converge weakly to the stationary distribution of the stochastic differential equation $$ dY_t=-h_0(Y_t)\,dt+Σ^{1/2}\,dB_t , $$ where $h_0$ is the limiting drift at the minimizer and $Σ$ denotes the asymptotic covariance. This recovers the Gaussian limit when $m=2$ and gives generally non-Gaussian stationary limits in the flat case $m>2$. Finally, we give corresponding results for coordinate-separable objectives with unequal flatness exponents.
Comments52 pages, 3 figures, 1 table