arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlexGrad:一种无根式的自适应步长方法

FlexGrad: a root-free approach to adaptive step sizes

Pablo Barros, Aaron Defazio, Vincent Guigues

arXiv 2610.09203首次发表:更新:

发表机构

Fundação Getulio Vargas; Meta(热图里奥·瓦加斯基金会; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FlexGrad是一种无根式自适应步长优化器,利用累积梯度和平方梯度缩放步长,在噪声与持续梯度场景间平衡,实现最优收敛,实验表现优异。

AI 中文摘要

我们引入了FlexGrad优化器,这是一种自适应梯度方法,解决了经典AdaGrad类方法的一个关键局限性:它们无法利用梯度之间的相关性。通过同时使用累积梯度和累积平方梯度来缩放步长,FlexGrad在噪声环境中保持了类似AdaGrad的行为,同时在梯度方向持续时允许显著更大的步长。该方法完全在线、无水平线,并且在优化过程中避免任何辅助搜索过程。对于凸Lipschitz目标,我们证明了FlexGrad在非光滑凸优化中达到了最优渐近收敛速率,并为全局和坐标-wise变体提供了保证。FlexGrad实现简单,并且与实际的梯度归一化方法紧密相关。在一系列凸和非凸问题上的实验表明,FlexGrad在各种优化设置中表现出竞争力。

英文摘要

We introduce the FlexGrad optimizer, an adaptive gradient method that addresses a key limitation of classical AdaGrad-type methods: their inability to exploit correlations between gradients. By scaling steps using both cumulative gradients and cumulative squared gradients, FlexGrad preserves AdaGrad-like behavior in noisy regimes while allowing substantially larger steps when gradient directions are persistent. The method is fully online, horizon-free, and avoids any auxiliary search procedure during optimization. For convex Lipschitz objectives, we prove that FlexGrad attains the optimal asymptotic convergence rate for nonsmooth convex optimization, with guarantees for both global and coordinate-wise variants. FlexGrad is straightforward to implement and is closely connected to practical gradient-normalized methods. Experiments on a range of convex and nonconvex problems show that FlexGrad performs competitively across a range of optimization settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑