发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文刻画了非凸-强凹极小极大优化的新鲜梯度复杂度,给出匹配上下界,并证明全局随机成本中线性条件数依赖的必要性及统计精炼成本的存在。
AI 中文摘要
我们刻画了光滑非凸-强凹极小极大优化的新鲜梯度预言复杂度,给出了匹配的上下界(至多相差对数因子)。设 $\Phi(x)=\max_y f(x,y)$,其中 $f$ 在无约束欧几里得域上关于 $y$ 是联合 $L$-光滑且 $\mu$-强凹的,并令 $\kappa=L/\mu$。每次查询返回一个新鲜的无偏联合梯度,其条件方差至多为 $\sigma^2$。给定初始原始间隙至多为 $\Delta$,对偶残差 $\\|\nabla_y f(x_0,y_0)\\|\le G$,在固定预算下,以至少 $2/3$ 的概率找到 $\\|\nabla\Phi(\widehat x)\\|\le\epsilon$ 的复杂度为 $\widetilde{\Theta}(\sqrt{\kappa}L\Delta/\epsilon^2+\kappa L\Delta\sigma^2/\epsilon^4+\kappa^2\sigma^2/\epsilon^2+\sqrt{\kappa}\log_+(G/(\epsilon\sqrt{\kappa})))$,其中 $\log_+u=\log\max\{1,u\}$。当 $\kappa$ 和 $L\Delta/\epsilon^2$ 超过通用常数时,该刻画在有限维数上一致成立,且被抑制的对数因子与 $G$ 无关。它确立了全局随机成本中线性条件数依赖的必要性,并识别出一个独立的统计精炼成本。一种近端方法将粗略的原始进展与最终的一次精炼分开,而下界适用于任意自适应随机算法。对偶初始化仅贡献一个加性对数成本,但移除其控制会消除所有有限维无维度复杂度界,即使使用精确梯度也是如此。
英文摘要
We characterize the fresh-gradient oracle complexity of smooth nonconvex-strongly-concave minimax optimization, with matching upper and lower bounds up to logarithmic factors. Let $Φ(x)=\max_y f(x,y)$, where $f$ is jointly $L$-smooth and $μ$-strongly concave in $y$ on unconstrained Euclidean domains, and set $κ=L/μ$. Each query returns a fresh unbiased joint gradient with conditional variance at most $σ^2$. Given initial primal gap at most $Δ$ and dual residual $\|\nabla_y f(x_0,y_0)\|\le G$, the fixed-budget complexity of finding $\|\nablaΦ(\widehat x)\|\leε$ with probability at least $2/3$ is $\widetildeΘ(\sqrtκLΔ/ε^2+κLΔσ^2/ε^4+κ^2σ^2/ε^2+\sqrtκ\log_+(G/(ε\sqrtκ)))$, where $\log_+u=\log\max\{1,u\}$. This characterization holds when $κ$ and $LΔ/ε^2$ exceed universal constants, uniformly over finite dimensions, and the suppressed logarithms are independent of $G$. It establishes the necessity of linear condition-number dependence in the global stochastic cost and identifies a separate statistical refinement cost. A proximal method separates coarse primal progress from one final refinement, while the lower bounds apply to arbitrary adaptive randomized algorithms. Dual initialization contributes only an additive logarithmic cost, yet removing its control eliminates every finite dimension-free complexity bound, even with exact gradients.