arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27383cs.LG

重尾噪声下Adam优化器的收敛行为

The Convergence Behavior of Adam under Heavy-Tailed Noise

发表机构密歇根州立大学
查看机构详情
  • Michigan State University(密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Yijiang Pang

首次发表
浏览论文内容

中文总结 AI 辅助

研究重尾噪声下Adam优化器的收敛性,推广在线到非凸转换框架分析Adam,明确其在重尾噪声下的收敛特性、次优性及域半径控制下的收敛速率提升情况。

中文摘要 AI 辅助

我们建立了普通向量形式Adam优化器在重尾随机噪声下的首个收敛保证。尽管已有若干Adam变体在有界方差非凸优化中达到最优迭代复杂度,但当随机梯度仅具有某p∈(1,2]的有界p阶中心矩时,其行为尚不明确,而该情形在现代深度学习中愈发常见。为填补这一空白,我们将近期的在线到非凸转换框架推广至适配重尾鞅差噪声。基于该广义框架,我们对Adam开展了无严格参数耦合的折扣后悔分析。结果表明,Adam在重尾噪声下收敛至(ρ,ε)-驻点,但呈现次优迭代复杂度及依赖p的收敛性,该次优性在有界方差情形(p=2)中依然存在。当已知域半径并用于控制在线学习器输出时(相关文献的标准设置),收敛速率提升至匹配最优复杂度。这些发现为Adam在重尾场景下的鲁棒性与局限性提供了新的理论见解。

英文摘要

We establish the first convergence guarantees for the plain vector-form Adam optimizer under heavy-tailed stochastic noise. While several Adam variants are known to achieve optimal iteration complexity in bounded-variance nonsmooth nonconvex optimization, little is understood about their behavior when stochastic gradients admit only a bounded $p$-th central moment for some $p \in (1,2]$, a setting increasingly observed in modern deep learning. To address this gap, we generalize the recent online-to-nonconvex conversion framework to accommodate heavy-tailed martingale-difference noise. Building on this generalized framework, we develop a discounted regret analysis for Adam, without restrictive parameter coupling. Our results show that Adam converges to $(ρ,ε)$-stationary points under heavy-tailed noise. However, it exhibits a suboptimal iteration complexity and $p$-dependent convergence, a suboptimality that persists even in the bounded-variance case ($p=2$). Specifically, the $ε$-dominant term in the iteration complexity for reaching in-expectation stationarity is $T=\mathrm{O}\left(Δρ^{1/2}(G+σ)^{\frac{5p}{3p-4}}ε^{-\left(\frac{5p}{3p-4}+\frac{3}{2}\right)}\right)$ for $p\in(\frac{4}{3},2]$, which simplifies to $T=\mathrm{O}(ε^{-13/2})$ when $p=2$. When the domain radius is known and used to control the online-learner output, a standard setup in related literature, the convergence rate improves to match the optimal complexity. In this case, the $ε$-dominant iteration complexity is $T=\mathrm{O}\left(Δρ^{1/2}(G+σ)^{\frac{p}{p-1}}ε^{-\left(\frac{p}{p-1}+\frac{3}{2}\right)}\right)$ for $p\in(1,2]$, which simplifies to $T=\mathrm{O}(ε^{-7/2})$ when $p=2$. These findings provide new theoretical insight into the robustness and limitations of Adam in heavy-tailed regimes.

补充信息

↑