极小极大问题的子空间方法
Subspace methods for min-max problems
- Universität Wien(维也纳大学)
- Brandenburg University of Technology(勃兰登堡工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出四组无雅可比子空间方法求解非线性单调方程,用于大规模机器学习极小极大问题,在单调性和Lipschitz条件下建立全局收敛及复杂度界,数值实验验证了鲁棒性和效率。
AI中文摘要:
本文针对非线性单调方程引入了四组子空间方法,并应用于大规模机器学习问题。这些方法使用共轭梯度型的无雅可比子空间({\tt JFS})方向,并结合固定步长或由Solodov和Svaiter投影方法生成的变步长。为了确保收敛性不依赖于子空间方向的具体代数形式,我们施加了一个角度条件以及一个控制有效搜索方向的显式缩放规则。在算子的单调性和Lipschitz连续性下,我们建立了线搜索和固定步长框架的全局收敛性,以及最佳迭代残差率$O(\ell^{-1/2})$。在局部误差界下,到解集的距离满足更快的衰减$o(\ell^{-1/2})$。如果算子连续可微且其雅可比矩阵在解处非奇异,则所需的局部误差界和局部孤立性成立,从而产生$R$-线性局部收敛。残差序列随后几何收敛,因此满足最后迭代率$o(\ell^{-1})$,而无需强单调性。我们还推导了迭代和残差评估的复杂度界:基线最佳迭代保证为$O(\varepsilon^{-2})$,局部线性区域为$O(\log(\varepsilon^{-1}))$,以及回溯残差评估的统一界。在额外的渐近假设下,所提出的{\tt JFS}方向和几种经典更新方向允许相关的乐观梯度下降-上升(\texttt{OGDA})型残差-记忆表示。数值实验展示了这些方法在代表性极小极大问题上的鲁棒性和效率。
英文摘要:
This paper introduces four groups of subspace methods for nonlinear monotone equations, with applications to large-scale machine learning problems. The methods use Jacobian-free subspace ({\tt JFS}) directions of conjugate-gradient type, combined with either fixed step sizes or variable step sizes generated by the projected method of Solodov and Svaiter. To ensure convergence independently of the specific algebraic form of the subspace directions, we impose an angle condition together with an explicit scaling rule controlling the effective search directions. Under monotonicity and Lipschitz continuity of the operator, we establish global convergence for both the line-search and fixed-step frameworks, as well as a best-iterate residual rate $O(\ell^{-1/2})$. Under a local error bound, the distance to the solution set satisfies the sharper decay $o(\ell^{-1/2})$. If the operator is continuously differentiable and its Jacobian is nonsingular at a solution, the required local error bound and local isolation follow, yielding $R$-linear local convergence. The residual sequence then converges geometrically and hence satisfies the last-iterate rate $o(\ell^{-1})$, without strong monotonicity. We also derive iteration and residual-evaluation complexity bounds: $O(\varepsilon^{-2})$ for the baseline best-iterate guarantee and $O(\log(\varepsilon^{-1}))$ in the local linear regime, together with a uniform bound on backtracking residual evaluations. Under additional asymptotic assumptions, the proposed {\tt JFS} directions and several classical update directions admit related optimistic gradient descent--ascent (\texttt{OGDA})-type residual--memory representations. Numerical experiments illustrate the robustness and efficiency of the methods on representative min--max problems.