如何使约束随机极小极大问题及其扩展中的梯度映射变小
How to Make the Gradient Mapping Small for Constrained Stochastic Min-Max Problems and Beyond
- University of British Columbia(不列颠哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究将约束凸凹极小极大问题的梯度映射复杂度从$\widetilde{O}(\varepsilon^{-4})$改进至$\widetilde{O}(\varepsilon^{-2})$,并推广至无界方差情形。
AI中文摘要:
我们研究了约束或正则化凸凹极小极大优化和随机单调变分不等式的随机一阶预言机复杂度。我们关注次优性以梯度映射(也称为前向-后向或自然残差)衡量的情况,这是一种将无约束问题的梯度范数推广的最优性概念。在此设置下,在标准无偏预言机访问和现在标准的方差假设下,使梯度映射范数小于$\varepsilon$的已知最佳复杂度为$\widetilde{O}(\varepsilon^{-4})$,而无约束情形下已建立近最优的$\widetilde{O}(\varepsilon^{-2})$。我们弥合了这一差距,将约束凸凹极小极大问题的梯度映射复杂度改进为$\widetilde{O}(\varepsilon^{-2})$。然后,我们通过使用Blum-Gladyshev假设,将相同的复杂度推广到无界方差的问题。
英文摘要:
We study the stochastic first-order oracle complexity for constrained or regularized convex-concave min-max optimization and stochastic monotone variational inequalities. We focus on the case when suboptimality is measured in terms of the gradient mapping, also known as, forward-backward or natural residual, an optimality notion that generalizes the gradient norm for unconstrained problems. In this setting, under standard unbiased oracle access with now-standard variance assumptions, the best-known complexity for making the norm of the gradient mapping less than $\varepsilon$ is $\widetilde{O}(\varepsilon^{-4})$, compared to the near-optimal $\widetilde{O}(\varepsilon^{-2})$ that is established in the unconstrained case. We bridge this gap to improve the gradient mapping complexity for constrained convex-concave min-max problems to $\widetilde{O}(\varepsilon^{-2})$. We then extend to prove the same complexity for problems without the bounded variance, by using the Blum-Gladyshev assumption.