AI 中文总结
本文为实用Muon(含Newton–Schulz迭代与解耦权重衰减)建立收敛性,证明其在随机非凸下以O(T^{-1/4})速率收敛,维度依赖优于AdamW。
AI 中文摘要
Muon正成为大规模神经网络训练中AdamW的有前景替代方案,然而对其实际实现的理论理解仍不完整,因为现有分析常常简化或省略两个关键组成部分:(i) 使用经验调优多项式系数$(3.4445,-4.7750,2.0315)$的实用Newton–Schulz迭代;(ii) 用于正则化的解耦权重衰减。在本文中,我们为实用Muon提供了优化解释并建立了收敛性,同时考虑了这两个组成部分。具体而言,我们将实用Muon解释为对原始损失的右预处理优化,并带有动态加权的$\ell_2$正则化项,该正则化项在接近平稳点时消失,从而优化目标保持为原始目标。然后,我们据我们所知,首次在随机非凸设置中为实用Muon建立了收敛保证,其收敛速率为$\mathcal{O}(T^{-1/4})$,以梯度的期望Frobenius范数衡量,将已知最佳AdamW收敛率的维度依赖性改善了$\sqrt{d}$倍,其中$T$为迭代次数,$d$为参数维度。实验进一步支持理论收敛结果。
英文摘要
Muon is emerging as a promising alternative to AdamW for large-scale neural network training, yet theoretical understanding of its practical implementation remains incomplete, as existing analyses often simplify or omit two key components: (i) practical Newton--Schulz iterations with empirically tuned polynomial coefficients $(3.4445,-4.7750,2.0315)$; and (ii) decoupled weight decay for regularization. In this paper, we provide an optimization interpretation and establish convergence for practical Muon, jointly accounting for both components. Specifically, we interpret practical Muon as right-preconditioned optimization of the original loss with a dynamic weighted $\ell_2$ regularizer that vanishes as stationarity is approached, so that the optimization target remains the original objective. We then establish, to our best knowledge, the first convergence guarantee for practical Muon in the stochastic nonconvex setting, with an $\mathcal{O}(T^{-1/4})$ convergence rate in terms of the expected Frobenius norm of the gradient, improving the dimension dependence of the best known AdamW's convergence rate by a factor of $\sqrt{d}$, where $T$ is the iteration horizon and $d$ is the parameter dimension. Experiments further support the theoretical convergence results.