发表机构
Meiji University(明治大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究普通SGD尤其是带动量的在重尾噪声下的表现,通过细化收敛结果及全面分析,揭示其在重尾噪声下收敛率低于裁剪或归一化变体,展现普通方法局限性,合成函数实验支持理论发现。
AI 中文摘要
随机梯度下降(SGD)是现代优化的基石。虽然其在重尾噪声下的性能常通过梯度裁剪或归一化等专门修改来解决,但本文研究更基本的问题:普通SGD,特别是带动量的SGD,在重尾噪声存在时如何表现?我们细化了普通SGD的现有收敛结果,更重要的是,首次对带动量的普通SGD在强凸、凸和非凸目标下进行了全面收敛分析,且未采用任何梯度控制机制。结果表明所得收敛率低于SGD裁剪或归一化变体的最优率,揭示了普通方法在重尾噪声下的固有局限性。理论发现得到了合成函数实验的支持。
英文摘要
Stochastic gradient descent (SGD) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundamental question: how does vanilla SGD, particularly with momentum, perform in the presence of heavy-tailed noise? In this paper, we refine existing convergence results for vanilla SGD and, more importantly, provide the first comprehensive convergence analysis of vanilla SGD with momentum for strongly convex, convex, and nonconvex objectives, without employing any gradient control mechanisms. Our results demonstrate that the obtained convergence rates are inferior to the optimal rates achieved by clipped or normalized variants of SGD, thereby revealing inherent limitations of vanilla methods under heavy-tailed noise. The theoretical findings are supported by experiments on synthetic functions.
CommentsAccepted at UAI2026