发表机构
UCLouvain, ICTEAM/INMA(天主教鲁汶大学 ICTEAM/INMA)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨光滑凸优化中梯度范数最小化的加速问题,反驳了逐点下界猜想,证明视界无关方法存在尖锐的无穷多视界下界,并揭示不同保证之间的不可兼得性。
AI 中文摘要
在光滑凸优化中,梯度范数是平稳性的直接可观测度量。加速最小化梯度范数的一阶方法比加速最小化函数值更为微妙。已知对于任何给定的有限视界,存在最优加速方法(如OGM-G,Kim & Fessler, 2021),但其系数显式依赖于视界的长度,即迭代次数。我们研究当停止视界对方法未知时(即视界无关方法),何种加速是可行的。Diakonikolas & Wang (2022) 猜想:对于任何非自适应、视界无关的线性跨度一阶方法,在每个视界N上,平方梯度范数存在Ω(N^{-1})的下界。我们通过展示一种在密度为一的视界集合上实现接近N^{-2}末次迭代保证的方法,反驳了这一逐点猜想。我们转而证明,对于无穷多个视界,Ω(N^{-1})下界必须成立。更精确地,若G_N(A)表示方法A在N次迭代后平方梯度范数的界,我们证明对于任何方法A,limsup_{N→∞} N G_N(A) ≥ 1/2。这个界是尖锐的:常数1/2恰好由Rotaru等人(2026)的视界无关梯度下降调度达到。此外,我们表明上述两种极端行为不能由同一方法实现:任何在迭代子序列上具有o(N^{-1})保证的方法A必须满足limsup_N N G_N(A)=∞。相比之下,最佳迄今为止输出享有均匀的O(N^{-2})保证。
英文摘要
In smooth convex optimization, the gradient norm is a directly observable measure of stationarity. Accelerating a first-order method that minimizes the gradient norm is known to be more delicate than accelerating the minimization of function values. Optimal accelerated methods such as OGM-G (Kim & Fessler, 2021) are known to exist for any prescribed finite horizon, but their coefficients depend explicitly on the length of that horizon, i.e. the number of iterations. We ask what kind of acceleration is feasible when the stopping horizon is unknown to the method, i.e. for horizon-independent methods. Diakonikolas & Wang (2022) conjectured that an $Ω(N^{-1})$ lower bound on the squared gradient norm holds at every horizon $N$ for any nonadaptive, horizon-independent linear-span first-order method. We disprove this pointwise conjecture by exhibiting a method that achieves near-$N^{-2}$ last-iterate guarantees on a density-one set of horizons. We show instead that an $Ω(N^{-1})$ lower bound must hold for infinitely many horizons. More precisely, if $\mathcal G_N(\mathcal A)$ denotes the bound on the squared gradient norm after $N$ iterations for a method $\mathcal A$, we prove that $\limsup_{N\to\infty}N \mathcal G_N(\mathcal A) \ge 1/2$ for any method $\mathcal A$. This bound is sharp: the constant $1/2$ is exactly attained by the horizon-independent gradient-descent schedule of Rotaru et al. (2026). In addition, we show that the above two extreme behaviors cannot be achieved by the same method: any method $\mathcal A$ with an $o(N^{-1})$ guarantee on a subsequence of iterates must satisfy $\limsup_N N \mathcal G_N(\mathcal A)=\infty$. In contrast, best-so-far output admits a uniform $O(N^{-2})$ guarantee.