arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新审视随机凸优化中的AdaGrad:最后迭代、高概率与下界

Revisiting AdaGrad in Stochastic Convex Optimization: Last Iterates, High Probability, and Lower Bounds

Weiming Ou, Xiao Wang

arXiv 2609.32729首次发表:更新:

发表机构

Shanghai University of Finance and Economics(上海财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文重新审视随机凸优化中的AdaGrad方法,证明其无通用最后迭代速率,并给出紧下界,同时表明有界迭代条件下可消除log T损失,达到O(1/√T)速率。

AI 中文摘要

AdaGrad和AdaGrad-Norm是广泛使用的自适应方法,但它们在随机凸优化中的精确行为仍未被充分理解。我们首先证明,即使在亚高斯噪声和有界迭代下,AdaGrad-Norm和AdaGrad也不具备任何通用的最后迭代速率。接着我们证明,仅有界方差条件太弱:即使迭代有界,也无法得到高概率的平均迭代速率;而若迭代无界,甚至可能无法保证期望意义上的收敛。此外,我们在亚高斯噪声下为AdaGrad-Norm和AdaGrad构造了紧的Ω(log T/√T)下界,表明现有平均迭代上界中的log T项是不可避免的。最后,我们证明一旦施加有界迭代条件,这个log T损失就会消失:在一般的ABC条件下,两种方法均达到O(1/√T)的速率。

英文摘要

AdaGrad and AdaGrad-Norm are widely used adaptive methods, but their precise behavior in stochastic convex optimization remains less understood. We first show that AdaGrad-Norm and AdaGrad do not admit any universal \textbf{last-iterate} rate, even under sub-Gaussian noise and bounded iterates. We then prove that bounded variance alone is too weak: even with bounded iterates, it cannot yield \textbf{high-probability average-iterate} rates, and without bounded iterates it may not even guarantee convergence in expectation. Besides, we construct tight $Ω(\log T/\sqrt T)$ lower bounds for both AdaGrad-Norm and AdaGrad under sub-Gaussian noise, showing that $\log T$ in existing average-iterate upper bounds is unavoidable. Finally, we show that this $\log T$ loss disappears once bounded-iterate condition is imposed: under a general ABC condition, both methods achieve rates ${O}(1/\sqrt{T})$.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑