发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对Muon优化器提出频谱靶向Muon,通过调整奇异值正交化阈值τ在归一化SGD与Muon间插值,经实验揭示Muon成功的机制及与AdamW的差异。
AI 中文摘要
Muon优化器会正交化每个更新矩阵,将其所有奇异值设为1,已被证明在训练大型语言模型时非常有效。然而,目前尚不清楚这种成功是源于放大了梯度下降忽略的小奇异方向,还是源于抑制了破坏训练的大的退化方向。我们提出了频谱靶向Muon,它仅正交化阈值τ以上或以下的奇异值,因此改变τ可在归一化SGD与Muon之间进行插值。该方法通过在移位Gram矩阵上用Newton-Schulz迭代计算的投影来隔离相关的奇异子空间,无需使用SVD。我们在CIFAR-10和NanoGPT speedruns上评估这些变体,跟踪梯度、更新和权重矩阵的有效秩,以及一个新指标——更新与权重矩阵等谱流形切空间的对齐度。我们发现三点结论:第一,动量的小奇异值并非噪声;正交化每个矩阵除少数最大奇异值外的所有部分,几乎与Muon效果相当,且仅触及动量的一小部分,而仅正交化顶部部分则效果远逊,尽管其包含几乎所有动量;在语言模型上,每个动量奇异值都远小于1,因此靶向正交化只能进行放大,Muon的优势在于将小到无法训练的方向调整为原始尺度后可训练。第二,缩小最大奇异值是保持参数谱平坦的原因,该操作成本低,且并非驱动损失的因素。第三,AdamW与Muon的主要区别在于构建结构的速度更慢,这解释了其启动更慢以及预热对AdamW有帮助但仅对Muon有害的原因。
英文摘要
The Muon optimizer orthogonalizes each update matrix, setting all of its singular values to one, and has proven highly effective for training large language models. It remains unclear, however, whether this success comes from amplifying small singular directions that gradient descent neglects or from suppressing large, degenerate directions that disrupt training. We introduce Spectrally Targeted Muon, which orthogonalizes only the singular values above or below a threshold $τ$, so that varying $τ$ interpolates between normalized SGD and Muon. It isolates the relevant singular subspaces with projections computed by Newton-Schulz iteration on a shifted Gram matrix, so no SVD is needed. We evaluate these variants on the CIFAR-10 and NanoGPT speedruns, tracking the effective rank of gradient, update, and weight matrices and a new metric, the alignment of updates with the tangent space of the weight matrix's isospectral manifold. We find three things. First, the small singular values of the momentum are not noise. Orthogonalizing everything except the few largest singular values of each matrix nearly matches Muon while touching only a small fraction of the momentum, whereas orthogonalizing only the top falls well short even though it holds almost all of it. On language models every momentum singular value is far below one, so targeted orthogonalization can only amplify, and Muon wins by making directions that are too small to train on at their raw scale trainable. Second, shrinking the largest singular values is what keeps the parameter spectrum flat; this is cheap, and it is not what drives the loss. Third, AdamW differs from Muon mainly in how slowly it builds structure, which explains its slower start and why warmup helps AdamW but only hurts Muon.