arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低维块鞍点问题无需精确内层求解

Saddle-Point Problems with a Low-Dimensional Block Do Not Need Accurate Inner Solves

Ivan Fomin, Alexander V. Gasnikov

arXiv 2609.32441首次发表:更新:

发表机构

Yandex; Innopolis University(Yandex; 因诺波利斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低维块鞍点问题,提出证书传输方法,利用切片下界模型作为先验,避免精确内层求解,显著降低对x的调用次数,并在多种设置下达到最优复杂度。

AI 中文摘要

许多学习问题将高维参数块与少量对抗变量或对偶变量耦合:例如在少数组上的最差组风险,或在少数约束下的学习。我们考虑 $\min_{x\in X}\max_{y\in Y} f(x,y)$,其中 $X\subseteq\mathbb{R}^n$ 是紧凸集且强凸,条件数为 $\kappa$,并分别统计对 $x$ 和 $y$ 的 oracle 调用次数。教科书方法在 $y$ 上运行割平面法,并用加速方法将每个内层问题求解到精度 $\varepsilon$;该方法需要 $O(m\log(1/\varepsilon))$ 次对 $y$ 的调用,但对 $x$ 需要 $O(m\sqrt{\kappa}\log^2(1/\varepsilon))$ 次调用。我们证明内层问题无需精确求解。我们的方法,证书传输,保持单个切片的强凸下界模型,并将其作为下一个切片上短加速运行的先验。由于在 $y$ 上是凹的,每次对 $y$ 的调用要么切割局部化器,要么将下界模型移动到两个切片的混合,并保证下界增加。对于 $m=1$,这给出了 $\varepsilon$-鞍点,对 $x$ 的调用次数为 $O(\sqrt{\kappa}\log(1/\varepsilon))$(加上初始化项),对 $y$ 的调用次数为 $O(\log(1/\varepsilon))$;即使仅假设 $f(x,\cdot)$ 是凹且 Lipschitz 的,这两个次数都是最优的。对于一般 $m$,若局部化器的质心可计算,该方法对 $x$ 需要 $O((m+\sqrt{m\kappa})\log(1/\varepsilon))$ 次调用,对 $y$ 需要最优的 $O(m\log(1/\varepsilon))$ 次调用。在 $y$ 强凹的情况下,两点加速方法使对 $x$ 的调用次数与 $m$ 无关。

英文摘要

Many learning problems couple a high-dimensional block of parameters with a handful of adversarial or dual variables: worst-group risk over a few groups, or learning under a few constraints. We consider $\min_{x\in X}\max_{y\in Y} f(x,y)$ with $X\subseteq\mathbb{R}^n$, a compact convex set and strongly convex with condition number $κ$, and we count the oracle calls for $x$ and for $y$ separately. The textbook approach runs a cutting-plane method in $y$ and solves every inner problem to accuracy $\varepsilon$ with an accelerated method; it needs $O(m\log(1/\varepsilon))$ calls for $y$ but $O(m\sqrtκ\log^2(1/\varepsilon))$ calls for $x$. We show that the inner problems need not be solved accurately. Our method, certificate transport, keeps a strongly convex lower model of a single slice and uses it as a prior for a short accelerated run on the next slice. By concavity in $y$, every call for $y$ then either cuts the localizer or moves the lower model to a mixture of the two slices with a certified increase of the lower bound. For $m=1$ this gives an $\varepsilon$-saddle point after $O(\sqrtκ\log(1/\varepsilon))$ calls for $x$, up to an initialization term, and $O(\log(1/\varepsilon))$ calls for $y$; both counts are optimal, even though $f(x,\cdot)$ is only assumed to be concave and Lipschitz. For general $m$, with centers of gravity of the localizer treated as computable, the method needs $O((m+\sqrt{mκ})\log(1/\varepsilon))$ calls for $x$ and the optimal $O(m\log(1/\varepsilon))$ calls for $y$. Under strong concavity in $y$, a two-point accelerated method makes the number of calls for $x$ independent of $m$.

Comments40 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑