arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaGrad算法在广义光滑凸优化中的迭代收敛性研究

On the Iterate Convergence of AdaGrad for Generalized Smooth Convex Optimization

Mathieu Besançon, Tung Quoc Le

arXiv 2608.00482首次发表:更新:

AI 中文总结

本文针对广义光滑凸优化问题,证明了三种AdaGrad变体在不同光滑性条件下的迭代收敛性,明确了其自适应性条件,并通过数值示例指出额外几何假设对序列收敛的必要性。

AI 中文摘要

我们证明了AdaGrad算法族优化凸可微目标函数时的序列收敛结果。具体而言,当目标函数为凸且局部Lipschitz光滑时,我们给出了三种主要AdaGrad变体(AdaNorm、AdaDiag、AdaFull)迭代收敛的充要条件,解决了文献中遗留的问题。我们利用这一一般结果研究了三种变体在广义$(L_0,L_1)$-光滑性条件下的表现,证明了在足够小的恒定步长下其序列收敛。此外,在所谓的$(L_0,L_1)$-多项式可修正光滑性假设下(该假设是$(L_0,L_1)$广义光滑性性质的松弛,适用于$L$-光滑函数或单变量多项式等多种函数类),我们证明这些AdaGrad变体在任意学习率下均能实现序列收敛。该结果给出了AdaGrad具备自适应性的条件,即无需根据实例调整参数。最后,我们提供了AdaGrad在凸和非凸函数上行为的数值示例,尤其通过实验构造了一个反例,表明仅光滑性不足以保证AdaGrad类算法的序列收敛,还需额外的几何假设(如本文中的凸性或Kurdyka-Łojasiewicz不等式)才能得到序列收敛结果。

英文摘要

We prove sequential convergence results for the AdaGrad algorithm family optimizing convex differentiable objectives. Specifically, we provide necessary and sufficient conditions for the convergence of iterates for the three main AdaGrad variants (AdaNorm, AdaDiag, AdaFull) when the objective is convex and locally Lipschitz-smooth, closing the question left open from the literature. We harness this general result to study the three variants under the generalized $(L_0,L_1)$-smoothness condition and show sequential convergence for sufficiently small constant step size. Moreover, under the so-called $(L_0,L_1)$-polynomially modifiable smoothness assumption, which is a relaxation of the $(L_0,L_1)$ generalized smoothness property and is satisfied by many function classes such as $L$-smooth functions or univariate polynomials, sequential convergence for these AdaGrad variants is proved for arbitrary learning rates. This result provides conditions under which AdaGrad presents adaptivity, i.e., does not require tuning the parameters based on the instance. Finally, we provide numerical illustrations of the behavior of AdaGrad on convex and nonconvex functions. In particular, we construct a counterexample empirically showing that smoothness alone is not sufficient for the sequential convergence of AdaGrad-type algorithms, and suggesting that additional geometric hypotheses (e.g., convexity as in this paper, or the Kurdyka-Łojasiewicz inequality) are indispensable for sequential convergence results.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑