凸-凹极小极大优化的近最优高阶预言复杂度
Near-Optimal Higher-Order Oracle Complexity for Convex--Concave Minimax Optimization
浏览论文内容
中文总结 AI 辅助
本研究证明凸-凹极小极大优化中,任意自适应确定性和随机化算法的高阶查询复杂度下界与张量算法类相同,匹配现有上界,达到近最优精度指数。
中文摘要 AI 辅助
对于光滑的凸-凹极小极大优化,Chen等人(2026)的高阶下界适用于具有规定正则化泰勒模型更新的受限张量算法类。我们为任意自适应确定性和随机化算法建立了相同的下界,在对数因子范围内匹配Zhang等人(2026)的上界。固定整数$p\ge 2$,设$L_p>0$为直径至多$D_Z>0$的紧凸乘积域上目标函数$p$阶导数的Lipschitz常数上界。每次可行查询返回目标值及直至$p$阶的所有导数。对于精度$\epsilon>0$,设$Q_{\mathrm{tan}}=L_pD_Z^p/\epsilon$(切向残差)和$Q_{\mathrm{gap}}=L_pD_Z^{p+1}/\epsilon$(鞍点间隙)。设$T_E^{\mathrm{det}}(\epsilon)$和$T_E^{\mathrm{rand}}(\epsilon)$为准则$E\in\{\mathrm{tan},\mathrm{gap}\}$的高维极小极大查询复杂度,其中随机化算法在每个实例上的成功概率至少为$2/3$。我们的下界和现有上界给出:对于足够大的$Q_E$,$c_pQ_E^{2/(3p-1)}\le T_E^{\mathrm{rand}}(\epsilon)\le T_E^{\mathrm{det}}(\epsilon)\le C_pQ_E^{2/(3p-1)}[1+\log(3+Q_E)]^{6(p-1)}$,其中$c_p,C_p>0$仅依赖于$p$。因此,相同的精度指数适用于超越张量更新规则的情形,甚至适用于随机化查询和任意可行输出。证明构造了一个标量凸-凹链,其具有完全平坦的门以隐藏完整的导数信息。直接乘积域误差见证和自适应转录论证建立了两种准则的下界。
英文摘要
For smooth convex--concave minimax optimization, the higher-order lower bound of Chen et al. (2026) applies to a restricted tensor-algorithm class with prescribed regularized Taylor-model updates. We establish the same bound for arbitrary adaptive deterministic and randomized algorithms, matching, up to logarithmic factors, the upper bound of Zhang et al. (2026). Fix an integer $p\ge 2$ and let $L_p>0$ bound the Lipschitz constant of the objective's $p$-th derivative on a compact convex product domain of diameter at most $D_Z>0$. Each feasible query returns the objective value and all derivatives through order $p$. For accuracy $ε>0$, set $Q_{\mathrm{tan}}=L_pD_Z^p/ε$ for tangent residual and $Q_{\mathrm{gap}}=L_pD_Z^{p+1}/ε$ for saddle gap. Let $T_E^{\mathrm{det}}(ε)$ and $T_E^{\mathrm{rand}}(ε)$ denote the high-dimensional minimax query complexities for criterion $E\in\{\mathrm{tan},\mathrm{gap}\}$, with randomized success probability at least $2/3$ on every instance. Our lower bounds and the existing upper bound give $c_pQ_E^{2/(3p-1)}\le T_E^{\mathrm{rand}}(ε)\le T_E^{\mathrm{det}}(ε)\le C_pQ_E^{2/(3p-1)}[1+\log(3+Q_E)]^{6(p-1)}$ for sufficiently large $Q_E$, where $c_p,C_p>0$ depend only on $p$. Thus the same accuracy exponent holds beyond tensor update rules, even for randomized queries and arbitrary feasible outputs. The proof constructs a scalar convex--concave chain with exactly flat gates that hide complete derivative information. Direct product-domain error witnesses and adaptive transcript arguments establish the lower bounds for both criteria.
发表机构
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。