arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01662math.OCcs.LGstat.ML

非凸-凹极小极大优化中带方差缩减的随机一阶算法的下界

Lower Bounds for Stochastic First-Order Algorithms with Variance Reduction in Nonconvex--Concave Minimax Optimization

Jiayi Song, Zi Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本文为非凸-凹极小极大优化中允许方差缩减的随机一阶算法建立了复杂度下界,量化了精度、对偶域半径和噪声的影响,并补充了强凹情形下的下界。

中文摘要 AI 辅助

我们在非凸-凹极小极大优化中为随机一阶算法建立了复杂度下界,允许算法使用方差缩减。我们的主要贡献是针对允许方差缩减的零尊重算法类给出一个下界,这超越了现有一些下界所施加的算法限制。我们考虑具有 $L$-Lipschitz 连续联合梯度、对偶域为欧几里得半径至多 $D_Y$ 的紧凸集、以及由对偶变量上最大化目标定义的原始值函数且初始次优性至多 $\Delta$ 的目标。目标精度 $\varepsilon$ 通过参数为 $1/(2L)$ 的约束原始值函数的 Moreau 包络的梯度范数来度量。在方差至多 $\sigma^2$ 且具有均方光滑性的无偏随机一阶预言机下,我们证明了下界 $\Omega\\\\!\left(L^2D_Y\Delta\varepsilon^{-3}+L^3D_Y^2\Delta\sigma^2\varepsilon^{-6}\right)$。该结果量化了精度、对偶域半径和预言机噪声的影响,即使允许方差缩减。我们还为非凸-强凹极小极大优化建立了互补下界。在对偶强凹参数 $\mu>0$ 和条件数 $\kappa:=L/\mu$ 下,我们在有界方差预言机模型下得到 $\Omega\\\\!\left(L\Delta\sqrt{\kappa}\\\\,\varepsilon^{-2}+L\Delta\kappa\sigma^2\varepsilon^{-4}\right)$。在常数 $\bar L$ 的额外均方光滑性条件下,我们得到 $\Omega\\\\!\left(L\Delta\sqrt{\kappa}\\\\,\varepsilon^{-2}+\Delta\bar L\sigma\kappa^{3/2}\varepsilon^{-3}\right)$。这些结果共同识别了凹和强凹体制下的复杂度障碍,其中主要的非凸-凹下界对使用方差缩减的算法仍然有效。

英文摘要

We establish complexity lower bounds for stochastic first-order algorithms in nonconvex--concave minimax optimization, allowing algorithms to use variance reduction. Our main contribution is a lower bound for a zero-respecting algorithm class that permits variance reduction, extending beyond the algorithmic restrictions imposed by some existing lower bounds. We consider objectives with an $L$-Lipschitz continuous joint gradient, a compact convex dual domain of Euclidean radius at most $D_Y$, and a primal value function, defined by maximizing the objective over the dual variable, with initial suboptimality at most $Δ$. The target accuracy $\varepsilon$ is measured by the gradient norm of the Moreau envelope of the constrained primal value function with parameter $1/(2L)$. Under an unbiased stochastic first-order oracle with variance at most $σ^2$ and mean-square smoothness, we prove the lower bound $Ω\!\left(L^2D_YΔ\varepsilon^{-3}+L^3D_Y^2Δσ^2\varepsilon^{-6}\right)$. This result quantifies the dependence on accuracy, dual-domain radius, and oracle noise even when variance reduction is allowed. We also establish complementary lower bounds for nonconvex--strongly-concave minimax optimization. With dual strong-concavity parameter $μ>0$ and condition number $κ:=L/μ$, we obtain $Ω\!\left(LΔ\sqrtκ\,\varepsilon^{-2}+LΔκσ^2\varepsilon^{-4}\right)$ under the bounded-variance oracle model. Under the additional mean-square smoothness condition with constant $\bar L$, we obtain $Ω\!\left(LΔ\sqrtκ\,\varepsilon^{-2}+Δ\bar Lσκ^{3/2}\varepsilon^{-3}\right)$. Together, these results identify complexity barriers across the concave and strongly concave regimes, with the main nonconvex--concave bound remaining valid for algorithms that use variance reduction.

发表机构

  • Shanghai University(上海大学)

机构由 AI 辅助整理,请以论文原文为准。

↑