发表机构
École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明,在重尾噪声下,裁剪的去中心化随机梯度下降(DSGD)能实现阶最优收敛率并具有线性加速,与归一化方法形成对比,凸显了裁剪在保留幅度信息方面的优势。
AI 中文摘要
重尾噪声在现代机器学习中已被广泛观察到,这促使了梯度裁剪和归一化等方法的使用。虽然这些方法在中心化设置中已被充分理解,但在去中心化设置中却知之甚少,因为对局部梯度施加非线性会同时影响优化和共识。最近关于去中心化非凸优化的研究在重尾噪声下研究了裁剪和归一化,其中裁剪产生次优的收敛率,而归一化需要局部动量或小批量才能收敛。这引出了一个问题:一个使用非线性的基线去中心化方法能否在重尾噪声下实现最优收敛率?我们以裁剪的去中心化随机梯度下降($\mathtt{DSGD}$)给出肯定回答。对于在具有有界$p$阶矩噪声($p \in (1,2]$)下的光滑非凸代价,我们证明裁剪的$\mathtt{DSGD}$在高概率和期望意义下都实现了阶最优收敛率。此外,我们建立了在智能体数量上的线性加速,据我们所知,这尚未在带裁剪的去中心化方法中得到证明。关键技术要素是对共识差距的尖锐分析,该分析利用了裁剪的结构,将网络效应归入高阶项。我们的结果凸显了去中心化设置中裁剪与归一化之间的重要区别:虽然归一化的$\mathtt{DSGD}$可能无法收敛,但裁剪保留了幅度信息,使$\mathtt{DSGD}$能够收敛且达到阶最优。数值实验验证了我们的理论。
英文摘要
Heavy-tailed noise has been widely observed in modern machine learning, motivating the use of methods like gradient clipping and normalization. While these methods are well understood in centralized settings, much less is known in decentralized ones, where applying a nonlinearity to local gradients affects both optimization and consensus. Recent works on decentralized non-convex optimization have studied both clipping and normalization under heavy-tailed noise, with clipping yielding suboptimal rates and normalization needing local momentum or mini-batches to converge. This raises the question: can a baseline decentralized method using a nonlinearity achieve optimal convergence rates under heavy-tailed noise? We answer affirmatively with clipped decentralized SGD ($\mathtt{DSGD}$). For smooth non-convex costs under bounded $p$-th moment noise, $p \in (1,2]$, we show that clipped $\mathtt{DSGD}$ achieves order-optimal rates both with high probability and in expectation. Moreover, we establish a linear speed-up in the number of agents, which, to our knowledge, has not been shown for decentralized methods with clipping. The key technical ingredient is a sharp analysis of the consensus gap that exploits the structure of clipping, relegating network effects to higher-order terms. Our results highlight an important distinction between clipping and normalization in decentralized settings: while normalized $\mathtt{DSGD}$ can fail to converge, clipping retains magnitude information, enabling $\mathtt{DSGD}$ to be convergent and order-optimal. Numerical experiments validate our theory.
Comments36 pages, 5 figures, 2 tables