频谱分配:为何Muon优于Adam,以及如何改进Muon
Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon
- University of Cambridge(剑桥大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过频谱分析揭示Muon优于Adam的机制,提出SAMuon及其简化版SAMuon-lite,在多规模modded-nanogpt模型上,二者均优于AdamW和Muon,且SAMuon可减少13.3%-24.0%训练令牌。
AI中文摘要:
诸如Muon之类的正交优化器相比Adam能大幅加速大语言模型的预训练,但其机制尚未被完全理解。我们通过对Transformer损失景观的样本外频谱探测分析来研究这一问题:在实际训练轨迹的检查点处,我们将每个动量缓冲分解为其奇异方向,并在保留数据上估计每个方向的损失最优步长。所得的频谱分布是各向异性的,但在批次、训练阶段、优化器及模型规模下均保持稳定:处于稳定边缘(Edge-of-Stability)运行的易变头部,其步长远小于可耐受的主体部分,而主体部分则允许大得多的步长。该分布提供了一个统一的频谱分配解释,说明为何Muon优于Adam,而Adam又优于SGD。这也揭示了Muon均匀缩放的一个局限:它仍未充分利用主体部分。基于这一发现,我们引入了频谱感知Muon(Spectral-Aware Muon, SAMuon),该方法将头部保持在Muon尺度,并利用静态频谱先验放大主体部分。我们提供两个变体:完整的SAMuon通过低秩随机奇异值分解(SVD)遵循测得的分布,而简化版SAMuon-lite则通过秩1幂迭代使用两级近似。两种方法均未在Muon之外添加持久优化器状态或显著的额外浮点运算(FLOPs),且在标准假设下,两种方法的理想精确白化版本均保留了Muon的渐近收敛率。在参数规模从1.24亿到10亿的“modded-nanogpt”模型上,两个变体在所有评估的模型规模和批次大小配置下,均优于调优后的AdamW和Muon(Scion实现)基线。SAMuon达到与Muon相同验证损失所需的训练令牌减少了13.3%至24.0%,而SAMuon-lite则保留了大部分此类增益,且几乎没有额外的时钟时间开销。
英文摘要:
Orthogonal optimisers such as Muon can substantially accelerate large language model pretraining relative to Adam, yet the mechanism remains incompletely understood. We investigate this through an out-of-sample spectral probing analysis of Transformer loss landscapes. At checkpoints along real training trajectories, we decompose each momentum buffer into its singular directions and estimate the loss-optimal step size along each direction on held-out data. The resulting spectral profile is anisotropic yet stable across batches and training stages, and consistent across the optimisers and model scales: a volatile head operating at the Edge-of-Stability supports a much smaller step size than the tolerant bulk, which permits substantially larger steps. This profile provides a unified spectral allocation account of why Muon outperforms Adam, which outperforms SGD. It also exposes a limitation of Muon's uniform scaling: it still underutilises the bulk. Guided by this finding, we introduce Spectral-Aware Muon (SAMuon), which holds the head at the Muon scale and amplifies the bulk using a static spectral prior. We provide two variants: the complete SAMuon follows the measured profile using a low-rank randomised SVD and the simplified SAMuon-lite uses a two-level approximation via rank-one power iteration. Neither method adds persistent optimiser state or notable extra FLOPs beyond Muon at scale, and the idealised exact-whitening versions of both retain Muon's asymptotic convergence rate under standard assumptions. Across "modded-nanogpt" models from 124M to 1B parameters, both variants outperform tuned AdamW and Muon (Scion implementation) baselines in all evaluated model-scale and batch-size configurations. SAMuon requires 13.3% to 24.0% fewer training tokens to reach the same validation loss as Muon, while SAMuon-lite retains most of this gain with near-zero wall-clock overhead.