发表机构
Dalian University of Technology; Tsinghua University(大连理工大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究重尾噪声下随机极小极大优化,证明无修改的SGDA在非凸强凹和非凸凹设置下可收敛,并开发无裁剪算法Stoc-TRGDAM和Stoc-TRGDmax实现最优精度依赖。
AI 中文摘要
随机极小极大优化因其在现代机器学习中的应用而受到越来越多的关注,而现有的理论研究主要依赖于随机梯度的有界方差假设。在重尾噪声下,随机梯度仅具有有限的$p$阶矩($p\in(1,2]$),通常认为梯度裁剪或归一化是保证收敛所必需的。在这项工作中,我们重新审视了重尾噪声下的随机极小极大优化,并对随机梯度下降上升(SGDA)进行了全面的理论研究。我们首先证明,在不修改更新规则的情况下,vanilla SGDA可以在重尾噪声下在非凸-强凹(NC-SC)和非凸-凹(NC-C)设置中收敛,建立了SGDA在这些情形下的首个收敛保证。除了无正则化问题,我们进一步研究了正则化随机极小极大优化,其中由于归一化与近端结构之间的不兼容性,直接将梯度归一化纳入近端更新是非平凡的。我们通过开发新的无裁剪算法(即Stoc-TRGDAM和Stoc-TRGDmax)克服了这一困难,这两种算法都能在不使用梯度裁剪的情况下实现对目标精度的最优依赖。
英文摘要
Stochastic min-max optimization has attracted increasing attention due to its applications in modern machine learning, while existing theoretical studies mainly rely on the bounded variance assumption for stochastic gradients. Under heavy-tailed noise, where stochastic gradients only possess a finite $p$-th moment for $p\in(1,2]$, gradient clipping or normalization is commonly believed to be necessary to guarantee convergence. In this work, we revisit stochastic min-max optimization under heavy-tailed noise and provide a comprehensive theoretical study of stochastic gradient descent ascent (SGDA). We first show that vanilla SGDA, without any modification to its update rule, can converge under heavy-tailed noise in both nonconvex-strongly-concave (NC-SC) and nonconvex-concave (NC-C) settings, establishing the first convergence guarantees for SGDA in these regimes. Beyond unregularized problems, we further investigate regularized stochastic min-max optimization, where directly incorporating gradient normalization into proximal updates is nontrivial due to the incompatibility between normalization and proximal structures. We overcome this difficulty by developing new clipping-free algorithms, i.e., Stoc-TRGDAM and Stoc-TRGDmax, and they both can achieve the optimal dependence on the target accuracy without using gradient clipping.