发表机构
Indian Institute of Technology Delhi; Robert Bosch GmbH(印度理工学院德里分校; 罗伯特·博世有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对 Muon 优化器,提出有限时间训练损失保证,证明在充分宽两层 ReLU 网络上,五步调优 Newton--Schulz 正交化与动量累积结合,以高概率达到任意目标损失,且无需精确正交化。
AI 中文摘要
现有的 Muon 收敛性分析要么假设精确正交化,要么分析经典的 Newton--Schulz 多项式,且仅保证平稳性,因此 Muon 的五步调优 Newton--Schulz 更新保留了哪些性质以及这是否足以达到指定的神经网络训练损失,仍未得到解决。我们建立了一个有限时间训练保证,同时考虑了正交化前的动量累积和调优的有限步更新。对于具有固定随机输出权重和正定极限神经正切核的充分宽的两层 ReLU 网络的全批量训练,我们证明 Muon 在初始化上以高概率达到任意目标经验平方损失 $\varepsilon>0$。对于每个动量参数 $\mu\in[0,1)$,一个与目标相关的常数学习率,与 $(1-\mu)\sqrt{\varepsilon}$ 成正比,产生一个击中时间界 $O((1-\mu)^{-1}\varepsilon^{-1/2})$,其他问题参数固定。充分宽度与目标精度和动量无关。分析表明,调优的 Newton--Schulz 映射在限制更新谱范数的同时,保持与动量缓冲区的对齐。初始化附近梯度变化的控制将该对齐转移到当前梯度,确保下降直到达到目标,而无需精确正交化。数值实验在低于充分理论阈值的宽度下支持这些机制:梯度更新对齐保持在分析参考之上,在固定的教师-学生数据集上,六种宽度和五种学生初始化的全部 30 次运行均达到目标损失,同时保持核正定性。
英文摘要
Existing convergence analyses of Muon either assume exact orthogonalization or analyze classical Newton--Schulz polynomials, and guarantee only stationarity, so it is unresolved what Muon's five tuned Newton--Schulz steps preserve and whether that suffices to reach a prescribed neural-network training loss. We establish a finite-time training guarantee that accounts for both momentum accumulation before orthogonalization and the tuned finite-step update. For full-batch training of a sufficiently wide two-layer ReLU network with fixed random output weights and a positive-definite limiting neural tangent kernel, we prove that Muon reaches any target empirical squared loss $\varepsilon>0$ with high probability over initialization. For every momentum parameter $μ\in[0,1)$, a target-dependent constant learning rate proportional to $(1-μ)\sqrt{\varepsilon}$ yields a hitting-time bound of $O((1-μ)^{-1}\varepsilon^{-1/2})$, with other problem parameters fixed. The sufficient width is independent of both target accuracy and momentum. The analysis shows that the tuned Newton--Schulz map preserves alignment with the momentum buffer while bounding the update's spectral norm. Control of gradient variation near initialization transfers this alignment to the current gradient, ensuring descent until the target is reached without requiring exact orthogonalization. Numerical experiments support these mechanisms at widths below the sufficient theoretical threshold: gradient-update alignment remains above the analytical reference, and all 30 runs across six widths and five student initializations on a fixed teacher-student dataset reach the target loss while maintaining kernel positivity.
Comments22 pages, 5 figures