Hölder光滑性下归一化梯度下降的末次迭代收敛速率
Last-Iterate Convergence Rate of Normalized Gradient Descent under Hölder Smoothness
浏览论文内容
中文总结 AI 辅助
本文研究归一化梯度下降在Hölder光滑凸目标下的末次迭代收敛速率,提出恒定步长含对数开销的上界,并证明线性递减步长可达到最优阶,无需先验参数知识。
中文摘要 AI 辅助
归一化梯度下降是一种被广泛研究的自适应优化方法。现有的大多数分析关注最佳迭代点或迭代点的加权平均值,而实际实现通常返回末次迭代点。本文研究了凸的、$(\nu,M_\nu)$-Hölder光滑目标下归一化梯度下降的末次迭代收敛性。对于恒定步长,我们建立了$\nmathcal{O}\bigl((\nlog^2(T)/T)^{(1+\nnu)/2}\bigr)$的上界,相对于已知的最佳迭代点和加权平均迭代点的$\nmathcal{O}\bigl(T^{-(1+\nnu)/2}\bigr)$保证,包含了一个对数开销。对于$\nnu = 0$,该开销已知是不可避免的。我们通过性能估计问题(PEP)的数值结果补充了这一分析,研究了光滑设置下的有限时域最坏情况行为,以及该对数开销是否反映了恒定步长归一化梯度下降的内在局限性。然后我们证明,线性递减的步长可以产生$\nmathcal{O}\bigl(T^{-(1+\nnu)/2}\bigr)$的末次迭代保证,与最佳迭代点/加权平均保证的阶数匹配,且无需知道$\nnu$和$M_\nnu$。
英文摘要
Normalized gradient descent is a widely studied adaptive optimization method. Most existing analyses focus on the best iterate or a weighted average of the iterates, whereas practical implementations typically return the last iterate. In this paper, we study the last-iterate convergence of normalized gradient descent for convex, $(ν,M_ν)$-Hölder-smooth objectives. For a constant stepsize, we establish an upper bound of $\mathcal{O}\bigl((\log^2(T)/T)^{(1+ν)/2}\bigr)$, which contains a logarithmic overhead relative to the known $\mathcal{O}\bigl(T^{-(1+ν)/2}\bigr)$ guarantees for the best and weighted-average iterates. For $ν= 0$, this overhead is known to be unavoidable. We complement this analysis with numerical results based on the performance estimation problem (PEP), investigating the finite-horizon worst-case behavior in the smooth setting and whether the logarithmic overhead reflects an intrinsic limitation of constant-step normalized gradient descent. We then show that a linearly decreasing stepsize yields a last-iterate guarantee of $\mathcal{O}\bigl(T^{-(1+ν)/2}\bigr)$, matching the order of the best-iterate/weighted-average guarantees without requiring knowledge of $ν$ and $M_ν$.
发表机构
- Toyota Motor Corporation(丰田汽车公司)
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。