arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

$\eta$-重尾梯度噪声下SGD的高概率保证

High-Probability Guarantees for SGD under $β$-Heavy-Tailed Gradient Noise

Qijun Tong, Masahiro Ikeda, Ryota Kawasumi

arXiv 2609.32195首次发表:更新:

发表机构

The University of Electro-Communications; The University of Osaka; Gunma University(电气通信大学; 大阪大学; 群马大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对SGD在重尾梯度噪声下的高概率保证,提出基于Orlicz空间Young函数的β-重尾噪声模型,建立集中不等式并给出优化与总体风险平稳性的高概率界,同时分析梯度裁剪情形。

AI 中文摘要

随机梯度下降(SGD)广泛用于训练机器学习模型,但对训练数据进行子采样会在其更新中引入噪声。因此,高概率保证的强度和适用性关键取决于梯度噪声的尾部如何建模。深度学习中重尾梯度噪声的报道促使我们放宽SGD高概率分析中常用的有界噪声和次高斯假设。我们使用Orlicz空间理论中的Young函数在统一框架中描述噪声尾部。我们通过采用一个保持所有多项式矩有限且允许比次Weibull更重的尾部(包括对数正态分布)的Young函数来建模SGD梯度噪声。由此产生的类别称为$\eta$-重尾,其中$\eta$控制尾部重度。我们为$\eta$-重尾噪声建立了集中不等式,并将其与沿SGD轨迹的经验梯度与总体梯度之差的均匀界相结合,在轨迹假设下为光滑非凸损失获得优化和总体风险平稳性的高概率界。这些界不限于特定的学习率衰减规则,并明确展示了噪声尾部和学习率调度的影响。在Polyak-Łojasiewicz条件下,我们限制了最后一次迭代的风险。我们还分析了在$\eta$-重尾噪声模型下使用梯度裁剪的SGD。

英文摘要

Stochastic gradient descent (SGD) is widely used to train machine learning models, but subsampling the training data introduces noise into its updates. The strength and applicability of high-probability guarantees therefore depend critically on how the tails of gradient noise are modeled. Reports of heavy-tailed gradient noise in deep learning motivate relaxing the bounded-noise and sub-Gaussian assumptions commonly used in high-probability analyses of SGD. We use Young functions from Orlicz space theory to describe noise tails in a common framework. We model SGD gradient noise by adopting a Young function that preserves the finiteness of all polynomial moments while allowing tails heavier than sub-Weibull, including lognormal distributions. The resulting class is called $β$-heavy-tailed, with $β$ controlling the tail heaviness. We establish concentration inequalities for $β$-heavy-tailed noise and combine them with a uniform bound on the difference between empirical and population gradients along the SGD trajectory to obtain high-probability bounds on optimization and population-risk stationarity for smooth nonconvex losses under trajectory assumptions. The bounds are not restricted to a particular learning-rate decay rule and make explicit the effects of noise tails and learning-rate schedules. Under the Polyak-Łojasiewicz condition, we bound the risk at the last iterate. We also analyze SGD with gradient clipping under the $β$-heavy-tailed noise model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑