arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33383math.OCcs.LG

局部线性最小化预言机其实是一种投影方法!

Local LMO is Secretly a Projection Method!

  • King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
  • KAUST(阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Peter Richtárik, Ammar Mahran

AI总结:

本文证明局部线性最小化预言机(Local LMO)在球半径不超过Polyak半径时等价于欧几里得投影,并属于更广的投影方法族;对Hölder连续梯度目标,该族方法达到通用最优速率,匹配Nesterov的非加速通用梯度方法。

AI中文摘要:

局部线性最小化预言机(Ferris 和 Zavriev,1996;arXiv:2605.08850),即局部 LMO,在不需计算投影的情况下求解约束凸问题:它最小化可行集与当前迭代点周围球体交集上的线性模型。我们证明,只要球半径不超过 Polyak 半径,局部 LMO 步就是当前迭代点到可行集与一个将迭代点与解集分离的半空间交集上的欧几里得投影;尽管如此,该投影仅通过线性预言机即可计算。我们证明局部 LMO 属于更广泛的投影方法族,该族可按局部化半空间的深度进行索引。对于具有 $\vartheta$-Hölder 连续梯度的目标函数,该族中任何半空间足够深的方法都能使其前 $K$ 次迭代中的最优值以通用速率 $\mathcal{O}(K^{-(1+\vartheta)/2})$ 达到最优,与 Nesterov(2015)的非加速通用梯度方法相匹配。在 Polyak 半径下运行时,当约束最优解也是无约束最优解($\\|\nabla f(x_\star)\\| = 0$)时,局部 LMO 达到相同速率;否则达到速率 $\mathcal{O}(K^{-1/(2-\vartheta)})$。

英文摘要:

The local linear minimization oracle (Ferris and Zavriev, 1996; arXiv:2605.08850), or Local LMO, solves constrained convex problems without having to compute a projection: it minimizes a linear model over the intersection of the feasible set with a ball around the current iterate. We show that, whenever the ball radius does not exceed the Polyak radius, the Local LMO step is the Euclidean projection of the current iterate onto the intersection of the feasible set with a half-space that separates the iterate from the solution set; this projection is nonetheless computable by a linear oracle alone. We demonstrate that Local LMO belongs to a broader family of projection methods which may be indexed by the depth of the localizing half-space. For objectives with $\vartheta$-Hölder continuous gradient, every method from this family whose half-space lies sufficiently deep drives the best of its first $K$ iterates to optimality at the universal rate $\mathcal{O}(K^{-(1+\vartheta)/2})$, matching the non-accelerated universal gradient method of Nesterov (2015). Run at the Polyak radius, Local LMO attains the same rate when the constrained optima are also unconstrained ($\|\nabla f(x_\star)\| = 0$), and the rate $\mathcal{O}(K^{-1/(2-\vartheta)})$ otherwise.

↑