AI 中文总结
针对高维原始变量与低维约束对偶变量的鞍点问题,提出将二阶更新限于对偶空间、一阶方法用于原始空间的解耦框架,结合牛顿法与Nesterov加速,实现$\mathcal{O}(L/T^2)$收敛速率且无需强对偶凹性。
AI 中文摘要
鞍点问题在现代机器学习中扮演着重要角色,包括鲁棒优化和算法公平性。我们研究结构化凸-凹极小极大问题,其中原始变量位于高维无约束空间 $\mathbb{R}^d$,而对偶变量被限制在低维约束空间 $\mathbb{R}^J$,且 $J\ll d$。尽管现有理论框架已确立加速率 $\mathcal{O}(L/T^2)$ 是可能的,但在高维情况下,高效求解相关子问题可能成为计算瓶颈。为解决此问题,我们引入一种解耦技术。通过将二阶更新限制在低维对偶空间,并在高维原始空间采用一阶方法,我们的框架将牛顿方法与Nesterov加速相结合。对于单纯形约束的对偶变量以及与自和谐障碍相容的正则化器,所提出的算法在 $T$ 次外部迭代后,在原始目标次优性方面达到 $\mathcal{O}(L/T^2)$ 的收敛速率,且无需强对偶凹性。每次外部迭代的额外线性代数成本为 $\widetilde{\mathcal{O}}(dJ^2+J^{3.5})$,同时还需一次分量损失及其雅可比矩阵的评估。
英文摘要
Saddle-point problems play an important role in modern machine learning, including robust optimization and algorithmic fairness. We investigate structured convex--concave minimax problems where the primal variable lies in a high-dimensional, unconstrained space $\mathbb{R}^d$, while the dual variable is confined to a constrained, low-dimensional space $\mathbb{R}^J$, with $J\ll d$. Although existing theoretical frameworks establish that an accelerated rate of $\mathcal{O}(L/T^2)$ is possible, solving the associated subproblems efficiently can become a computational bottleneck in high dimensions. To address this, we introduce a decoupling technique. By restricting second-order updates to the low-dimensional dual space and employing first-order methods in the high-dimensional primal space, our framework combines Newton methods with Nesterov acceleration. For simplex-constrained dual variables and regularizers compatible with a self-concordant barrier, the resulting algorithm achieves a convergence rate of $\mathcal{O}(L/T^2)$ in primal objective suboptimality after $T$ outer iterations, without strong dual concavity. The additional linear-algebra cost per outer iteration is $\widetilde{\mathcal{O}}(dJ^2+J^{3.5})$, alongside one evaluation of the component losses and their Jacobian.