发表机构
MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究光滑凸优化中预定非负步长对梯度下降加速的极限,证明非任意时刻与任意时刻下的下界,结合已有上界,确定两种设置下的最优多项式收敛指数。
AI 中文摘要
我们研究了在光滑凸优化中,通过预定的非负步长,梯度下降(GD)能被加速到何种程度。记 $p_{\mathrm{sil}}=\log_2(1+\sqrt{2})$,我们证明了一个非任意时刻的下界 $\Omega\left(n^{-p_{\mathrm{sil}}-O(\sqrt{\log\log n/\log n})}\right)$。在任意时刻设置中,每个无限非负调度都有无穷多个时间范围,其误差为 $\Omega\left(n^{-\frac{2p_{\mathrm{sil}}}{1+p_{\mathrm{sil}}}-O(\sqrt{\log\log n/\log n})}\right)$。结合 silver 调度的上界 [Altschuler 和 Parrilo, 2025] 以及任意时刻上界 [Zhang 等人, 2025],我们的结果确定了两种设置下的最优多项式收敛指数。
英文摘要
We study how far gradient descent (GD) can be accelerated by predetermined stepsizes in smooth convex optimization. Writing $p_{\mathrm{sil}}=\log_2(1+\sqrt{2})$, we prove an $Ω\left(n^{-p_{\mathrm{sil}}-O(\sqrt{\log\log n/\log n})}\right)$ non-anytime lower bound. In the anytime setting, every infinite schedule has infinitely many horizons with error $Ω\left(n^{-\frac{2p_{\mathrm{sil}}}{1+p_{\mathrm{sil}}}-O(\sqrt{\log\log n/\log n})}\right)$. Together with the silver-schedule upper bound [Altschuler and Parrilo, 2025] and the anytime upper bound [Zhang et al., 2025], our results determine the optimal polynomial convergence exponents in both settings.
Comments35 pages, 4 figure