Silver 步长对梯度下降(几乎)最优:强凸情形
Silver Rate Is (Almost) Optimal for Gradient Descent: The Strongly Convex Case
浏览论文内容
中文总结 AI 辅助
本文证明在光滑强凸函数上,梯度下降采用 Silver 步长调度时,其迭代复杂度下界与已知上界匹配,从而表明 Silver 步长几乎最优。
中文摘要 AI 辅助
我们研究在光滑强凸函数上使用预定非负步长的梯度下降法。设 $p_{\mathrm{sil}}=\log_2(1+\sqrt2)$ 且 $\kappa$ 为条件数。我们证明了迭代下界 $\Omega\left(\kappa^{\frac{1}{p_{\mathrm{sil}}}-o(1)}\log\frac1\delta\right)$ 对相对平方距离和相对函数误差均成立,且对 $0<\delta<1$ 和足够大的 $\kappa$ 一致成立。这与 [Altschuler 和 Parrilo, 2025] 中建立的 Silver 步长调度的 $\kappa$ 多项式指数相匹配。
英文摘要
We study gradient descent with predetermined nonnegative stepsizes on smooth strongly convex functions. Let $p_{\mathrm{sil}}=\log_2(1+\sqrt2)$ and $κ$ be the condition number. We prove the iteration lower bound $Ω\left(κ^{\frac{1}{p_{\mathrm{sil}}}-o(1)}\log\frac1δ\right)$ for both relative squared distance and relative function error, uniformly over $0<δ<1$ and sufficiently large $κ$. This matches the polynomial exponent of $κ$ for the Silver stepsize schedule established in [Altschuler and Parrilo, 2025].
发表机构
- MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。