发表机构
University of Waterloo(滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明带裁剪和加性高斯噪声的随机梯度下降在光滑性和噪声有界假设下几乎必然收敛,并推广至动量变体,为裁剪方法的路径行为提供理论基础。
AI 中文摘要
带梯度裁剪和加性噪声的随机梯度下降(SGD)已成为训练机器学习模型的标准技术,特别是在需要鲁棒性或隐私保证的应用中。然而,裁剪在随机梯度中引入了偏差,而加性噪声引入了额外的方差,使得单个优化轨迹的长期行为难以刻画。在这项工作中,我们证明了在光滑性和随机梯度噪声一致有界的假设下,只要步长满足一些标准的衰减条件,带裁剪和加性高斯噪声的SGD(SGD-CN)几乎必然(a.s.)收敛。我们的分析扩展到动量变体,如随机重球和Nesterov加速梯度,在这些变体中,我们展示了仔细的能量构造能产生类似的保证。这些结果为理解裁剪随机梯度方法的路径行为提供了更强的理论基础,并表明尽管裁剪和扰动引入了偏差和噪声,该算法在凸和非凸区域中仍然保持稳定。
英文摘要
Stochastic gradient descent (SGD) with gradient clipping and additive noise has become a standard technique for training machine learning models, particularly in applications requiring robustness or privacy guarantees. However, clipping introduces a bias in stochastic gradients, while additive noise introduces additional variance, making the long-run behaviour of individual optimization trajectories difficult to characterize. In this work, we prove that SGD with clipping and additive Gaussian noise (SGD-CN) converges almost surely (a.s.) under smoothness and uniformly bounded stochastic-gradient noise assumptions, provided the step sizes satisfy some standard decaying conditions. Our analysis extends to momentum variants such as the stochastic heavy ball and Nesterov's accelerated gradient, where we show that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for understanding the pathwise behaviour of clipped stochastic gradient methods and suggest that, despite the bias and noise introduced by clipping and perturbation, the algorithm remains stable in both convex and nonconvex regimes.
Journal refIEEE Conference on Decision and Control (CDC), 2026