arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11056math.OC

约束凸-凹极小极大问题单环随机方法的最后迭代收敛性

Last-Iterate Convergence of Single-Loop Stochastic Methods for Constrained Convex-Concave Minimax Problems

Taoli Zheng, Jiajin Li, Anthony Man-Cho So

首次发表
浏览论文内容

中文总结 AI 辅助

研究约束凸-凹极小极大问题单环随机方法的最后迭代收敛性,引入扰动框架正则化问题,得到PS-EG和PS-OGDA方法。在已知或未知优化视界时分别给出收敛速率保证,无约束时PS-EG在梯度范数上有更优随时收敛速率。

中文摘要 AI 辅助

本文研究了在标准有界方差随机预言机下,用于约束光滑凸-凹极小极大优化的随机一阶方法的最后迭代收敛性。一个基本挑战是,即使对于简单的双线性问题,普通随机外梯度(S-EG)和随机乐观梯度下降-上升(S-OGDA)的最后迭代在存在随机梯度噪声时可能无法收敛。为克服这一困难,引入了一个简单的扰动框架,将原始凸-凹问题正则化为强凸-强凹问题。将S-EG和S-OGDA应用于扰动问题产生了两种简单的单环方法,即扰动S-EG(PS-EG)和扰动S-OGDA(PS-OGDA)。通过首先推导扰动问题到鞍点的平方距离的收敛性,然后将此估计转化为对受限原始-对偶间隙的保证,建立了最后迭代收敛性。基于此框架,建立了两种收敛保证。当优化视界先验已知时,PS-EG和PS-OGDA对于受限原始-对偶间隙都实现了\(\mathcal{O}(T^{-1/4})\)的最后迭代收敛速率,这与紧致可行域上的标准原始-对偶间隙一致。当优化视界未知时,基于递减扰动和递减步长开发了一种随时变体。对于一般的闭凸可行集,PS-EG和PS-OGDA对于受限原始-对偶间隙都实现了\(\mathcal{O}(T^{-1/5})\)的最后迭代收敛速率。此外,在无约束设置中,PS-EG在梯度范数方面允许更尖锐的\(\mathcal{O}(T^{-1/4})\)随时收敛速率。

英文摘要

In this paper, we study last-iterate convergence of stochastic first-order methods for constrained smooth convex--concave minimax optimization under the standard bounded-variance stochastic oracle. A fundamental challenge is that the last iterates of vanilla stochastic extragradient (S-EG) and stochastic optimistic gradient descent--ascent (S-OGDA) may fail to converge in the presence of stochastic gradient noise, even for simple bilinear problems. To overcome this difficulty, we introduce a simple perturbation framework that regularizes the original convex--concave problem into a strongly convex--strongly concave one. Applying S-EG and S-OGDA to the perturbed problem yields two simple single-loop methods, referred to as perturbed S-EG (PS-EG) and perturbed S-OGDA (PS-OGDA). We establish last-iterate convergence by first deriving convergence in terms of the squared distance to the saddle point of the perturbed problem and then translating this estimate into guarantees for the restricted primal--dual gap. Based on this framework, we establish two types of convergence guarantees. When the optimization horizon is known \emph{a priori}, both PS-EG and PS-OGDA achieve an $\mathcal{O}(T^{-1/4})$ last-iterate convergence rate for the restricted primal--dual gap, which coincides with the standard primal--dual gap on compact feasible domains. When the optimization horizon is unknown, we develop an anytime variant based on diminishing perturbations and diminishing stepsizes. For general closed convex feasible sets, both PS-EG and PS-OGDA achieve an $\mathcal{O}(T^{-1/5})$ last-iterate convergence rate for the restricted primal--dual gap. Furthermore, in the unconstrained setting, PS-EG admits a sharper $\mathcal{O}(T^{-1/4})$ anytime convergence rate in terms of the gradient norm.

补充信息

↑