基于步长的梯度下降加速的一个下界
A lower bound for stepsize-based acceleration of gradient descent
浏览论文内容
中文总结 AI 辅助
该研究针对光滑凸优化中仅靠步长调度加速的普通梯度下降,证明其最后一次迭代收敛速度的下界为$Ω(T^{-1.9319})$,说明仅步长调整无法让GD达到最优$O(T^{-2})$收敛速度,证明由GPT-5.6 Sol Pro在作者指导下完成。
中文摘要 AI 辅助
近期研究表明,在光滑凸优化问题中,仅通过精心设计的步长调度,无需借助动量或其他算法改进,普通梯度下降的收敛速度就能从教科书级的$O(T^{-1})$(其中$T$表示迭代次数)提升至$O\big(T^{-\log_2(1+\sqrt{2})}\big)$。尽管取得了这一进展,但除了通用一阶方法的经典$Ω(T^{-2})$基准之外,人们对这类方法的下界知之甚少。本研究针对采用预定非负步长调度的梯度下降的最后一次迭代收敛速度,提出了新的下界$Ω(T^{-1.9319})$。该结果提供了严格证据,表明仅靠步长调度无法将普通GD加速至最优的$O(T^{-2})$收敛速度。证明过程由GPT-5.6 Sol Pro在作者指导下完成。
英文摘要
Recent work has shown that, for smooth convex optimization, plain gradient descent can be accelerated from its textbook convergence rate of $O(T^{-1})$ (where $T$ denotes the number of iterations) to $O\big(T^{-\log_2(1+\sqrt{2})}\big)$ using carefully designed stepsize schedules alone, without resorting to momentum or other algorithmic modifications. Despite this progress, however, little was known about lower bounds for such methods beyond the classical $Ω(T^{-2})$ benchmark for general first-order methods. In this work, we present a new lower bound of $Ω(T^{-1.9319})$ for the last-iterate convergence rate of gradient descent with predetermined nonnegative stepsize schedules. This result provides rigorous evidence that stepsize schedules alone cannot accelerate plain GD to the optimal $O(T^{-2})$ convergence rate. The proof was developed by GPT-5.6 Sol Pro under the authors' guidance.
发表机构
- Tsinghua University(清华大学)
- University of Pennsylvania(宾夕法尼亚大学)
- Wharton School(沃顿商学院)
机构由 AI 辅助整理,请以论文原文为准。