arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30276cs.LGcs.SYeess.SY

为什么裁剪对AdaGrad至关重要?——广义光滑性下的高概率理论

Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness

  • Robert Bosch Center for Cyber Physical Systems, IISc Bengaluru(罗伯特博世网络物理系统中心,印度科学研究所班加罗尔)
  • Undergraduate Program, IISc Bengaluru(印度科学研究所班加罗尔本科项目)
  • Department of Computational Data Sciences, IISc Bengaluru(印度科学研究所班加罗尔计算数据科学系)
  • Department of Computer Science and Automation, IISc Bengaluru(印度科学研究所班加罗尔计算机科学与自动化系)
  • Tata Consultancy Services Research, Mumbai(塔塔咨询服务研究院,孟买)
  • System and Control Engineering, IIT Bombay(印度理工学院孟买分校系统与控制工程系)
  • Centre for Infrastructure, Sustainable Transportation and Urban Planning, IISc Bengaluru(印度科学研究所班加罗尔基础设施、可持续交通与城市规划中心)

机构由 AI 辅助整理,请以论文原文为准。

Alokendu Mazumder, Ayaan Mohd, Harshit Rawat, Arnab Roy, Mayank Baranwal, Punit Rathore

AI总结:

本文证明未裁剪AdaGrad在重尾噪声下会因各向异性失准而失效,而裁剪可修复该问题,并提供有限时域高概率收敛保证及O(ε^{-2})复杂度,表明裁剪是自适应几何的结构稳定器。

AI中文摘要:

我们在广义光滑性和有界方差的重尾噪声下,分析了原始的同步坐标AdaGrad算法。在该设定中,局部曲率可能随梯度范数次二次增长,且随机梯度仅假设具有有界条件二阶矩。我们证明,未裁剪的AdaGrad可能变得“各向异性失准”:在重尾噪声下,自适应分母可能学习到罕见噪声冲击的几何特征,而非目标函数的局部曲率,导致持续的方向性扭曲,阻碍有限时域内的欧几里得进展。随后,我们证明裁剪能修复这一失效模式。我们的主要结果是对原始非滞后AdaGrad更新的有限时域高概率保证,得到$\frac1T\sum_{t=0}^{T-1}\\|\nabla f(x_t)\\|^2=\mathcal{O}\left(\frac{d\big(\sqrt{\log T} + \log \frac{1}{\delta}\big)}{\sqrt{T}}\right)$,从而具有$\widetilde{\mathcal O}(\varepsilon^{-2})$复杂度。这表明,对于重尾噪声下的AdaGrad,裁剪是自适应几何的结构稳定器,而不仅仅是鲁棒性启发式方法。

英文摘要:

We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to have only bounded conditional second moments. We show that unclipped AdaGrad can become \emph{anisotropically miscalibrated}: under heavy-tailed noise, the adaptive denominator can learn the geometry of rare noise shocks rather than the local curvature of the objective, leading to a persistent directional distortion that blocks finite-horizon Euclidean progress. We then prove that clipping repairs this failure mode. Our main result is a finite-horizon high-probability guarantee for the original non-lagged AdaGrad update, yielding $\frac1T\sum_{t=0}^{T-1}\|\nabla f(x_t)\|^2=\mathcal{O}\left(\frac{d\big(\sqrt{\log T} + \log \frac{1}δ\big)}{\sqrt{T}}\right),$ and hence $\widetilde{\mathcal O}(\varepsilon^{-2})$ complexity. This shows that, for AdaGrad under heavy-tailed noise, clipping is a structural stabilizer of the adaptive geometry rather than merely a robustness heuristic.

↑