arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17973math.OCcs.LGstat.ML

匹配多循环复杂度与单循环:非凸-凹极小极大优化中的最优优化平稳性和最佳已知博弈平稳性

Matching Multi-Loop Complexities with a Single Loop: Optimal Optimization Stationarity and Best-Known Game Stationarity in Nonconvex--Concave Minimax Optimization

Minghao Zhang, Zi Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出单循环投影阻尼外梯度算法,在非凸-凹极小极大优化中同时达到优化与博弈平稳性的最佳已知复杂度,并证明预热版本在优化平稳性上最优。

中文摘要 AI 辅助

我们为光滑非凸-凹极小极大优化引入了一种新的单循环算法框架。由此产生的投影阻尼外梯度方法结合了投影外梯度更新、对偶动量和移动近端中心。在优化平稳性和博弈平稳性两种准则下,我们的方法在单循环一阶方法中实现了最佳已知复杂度。对于优化平稳性,我们的方法实现了梯度复杂度 $O(L^2D_Y\bar\Delta_0\varepsilon^{-3})$,其中 $L$ 是梯度Lipschitz常数,$D_Y$ 限制对偶可行集的直径,$\bar\Delta_0$ 是涉及值函数间隙和初始梯度的初始化量。此外,通过加入固定中心预热阶段,复杂度可以改进为 $O(L^2D_Y\Delta_\phi\varepsilon^{-3})$,加上一个加性的低阶成本,其中 $\Delta_\phi:=\phi(x_0)-\inf_x\phi(x)$。我们进一步建立了投影零尊重一阶方法在优化平稳性上的下界 $\Omega(L^2D_Y\Delta_\phi\varepsilon^{-3})$。该下界证明了我们算法的预热版本在该预言机类中对于优化平稳性在常数因子内是最优的。对于博弈平稳性,我们的方法实现了 $\mathcal{O}\\!(L^{3/2}D_Y^{1/2}\Delta_\phi\varepsilon^{-5/2})$ 梯度复杂度。这与多循环一阶方法的最佳已知复杂度相匹配,从而以单循环算法结构建立了相同的复杂度。在对偶强凹性下,所提出的框架对于两种平稳性准则实现了 $O\\!\sqrt{\kappa}\\,L\Delta_\phi\varepsilon^{-2}$ 的主要复杂度,其中 $\kappa=L/\mu$ 是对偶条件数,加上加性的初始化成本。$\varepsilon^{-2}$ 的精度依赖在固定正则性和初始化界限下是最优的。

英文摘要

We introduce a new single-loop algorithmic framework for smooth nonconvex--concave minimax optimization. The resulting projected damped extragradient method combines projected extragradient updates, dual momentum, and a moving proximal center. Under both the optimization-stationarity and game-stationarity criteria, our method achieves the best-known complexity among single-loop first-order methods. For optimization stationarity, our method achieves a gradient complexity of $O(L^2D_Y\barΔ_0\varepsilon^{-3})$, where $L$ is the gradient Lipschitz constant, $D_Y$ bounds the diameter of the dual feasible set, and $\barΔ_0$ is an initialization quantity involving the value-function gap and the initial gradients. Moreover, by incorporating a fixed-center warm-up phase, the complexity can be improved to $O(L^2D_YΔ_ϕ\varepsilon^{-3})$, up to an additive lower-order cost, where $Δ_ϕ:=ϕ(x_0)-\inf_xϕ(x)$. We further establish a lower bound of $Ω(L^2D_YΔ_ϕ\varepsilon^{-3})$ for optimization stationarity over projected zero-respecting first-order methods. This lower bound proves that the warm-started version of our algorithm is optimal up to a constant factor for optimization stationarity within this oracle class. For game stationarity, our method achieves $\mathcal{O}\!(L^{3/2}D_Y^{1/2}Δ_ϕ\varepsilon^{-5/2})$ gradient complexity. This matches the best-known complexity of multi-loop first-order methods, thereby establishing the same complexity with a single-loop algorithmic structure. Under dual strong concavity, the proposed framework achieves $O\!(\sqrtκ\,LΔ_ϕ\varepsilon^{-2})$ leading complexity for both stationarity criteria, where $κ=L/μ$ is the dual condition number, up to an additive initialization cost. The $\varepsilon^{-2}$ accuracy dependence is optimal under fixed regularity and initialization bounds.

发表机构

  • Shanghai University(上海大学)

机构由 AI 辅助整理,请以论文原文为准。

↑