Musec:用于稳定Muon型训练的动量谱裁剪
Musec: MomentUm SpEctral Clipping for Stable Muon-type Training
浏览论文内容
中文总结 AI 辅助
针对Muon优化器训练不稳定问题,提出动量谱裁剪(Musec)及高效Soft Musec实现,提供首个非凸非光滑收敛保证,并在多种设置下提升稳定性。
中文摘要 AI 辅助
Muon已成为大型语言模型训练中一种高效的优化器,与广泛采用的Adam和AdamW优化器相比,通常能实现更优的收敛性和性能。然而,Muon由于其谱平坦化特性,容易导致训练不稳定,表现为损失尖峰和模型权重的无界增长。现有方法主要依赖于权重或注意力logit裁剪,这些方法需要针对特定架构进行修改,且不能直接解决所有模型组件中的不稳定性。我们提出了动量谱裁剪(Musec),用谱裁剪取代Muon的谱平坦化:不是将动量矩阵的所有奇异值设置为近似一,而是裁剪超过阈值的奇异值,同时保留动量的底层谱结构。我们的策略提供了一种优化器级别的、架构无关的机制来稳定Muon训练。我们进一步开发了Soft Musec,一种高效实现,使用由耦合的Newton-Schulz迭代近似的平滑谱饱和函数。理论上,我们在非凸非光滑随机优化中为Musec建立了收敛保证。据我们所知,这是Muon型方法在非凸非光滑设置中的首个收敛保证。我们提供的实证研究表明,Soft Musec在广泛的学习率和模型规模范围内,持续优于现有Muon变体,提高了训练稳定性。值得注意的是,Soft Musec在现有Muon变体发散的设置中保持稳定,同时在良好调优的配置下匹配其性能。
英文摘要
Muon has emerged as a highly effective optimizer for large language model training, often achieving superior convergence and performance compared with the widely adopted Adam and AdamW optimizers. Nevertheless, Muon is prone to training instability due to its spectral flattening, manifested by loss spikes and unbounded growth of model weights. Existing approaches primarily rely on weight or attention-logit clipping, which require architecture-specific modifications and do not directly address instability across all model components. We propose MomentUm SpEctral Clipping (Musec), which replaces Muon's spectral flattening with spectral clipping: rather than setting all singular values of the momentum matrix to approximately one, Musec clips singular values that exceed a threshold while preserving the underlying spectral structure of the momentum. Our strategy provides an optimizer-level, architecture-agnostic mechanism for stabilizing Muon training. Theoretically, we establish convergence guarantees for Musec in nonconvex nonsmooth stochastic optimization. Practically, we develop Soft Musec, an efficient implementation that uses a smooth spectral saturation function approximated by coupled Newton-Schulz iterations. Empirically, Soft Musec consistently improves training stability over existing Muon variants across a wide range of learning rates and model sizes, remaining stable in settings where existing Muon variants diverge while matching their performance under well-tuned configurations.
发表机构
- National University of Singapore(新加坡国立大学)
- Wonders Information Co., Ltd.(万达信息股份有限公司)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。