用于μ子的各向同性保持谱帽:理论与三个案例研究
An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies
浏览论文内容
中文总结 AI 辅助
研究μ子和相关矩阵符号优化器对权重矩阵的影响,提出基于尺度不变性假设的统一框架及轻量级“谱帽”,通过三个案例研究表明谱帽可增加各向同性并防止失败,验证损失基本不变,结果为初步的。
中文摘要 AI 辅助
μ子和相关矩阵符号优化器越来越多地用于预训练大型语言模型,但它们对单个权重矩阵内部几何结构的影响尚不清楚。本初步报告提出了一个统一框架,基于一个理想化假设——权重重新缩放时损失的精确尺度不变性,这在高度归一化的网络中近似成立。在此假设下,普通随机梯度下降(SGD)在更新大小上有一个内置的1/||W||制动,而μ子的矩阵符号步移除了该制动,因此弗罗贝尼乌斯范数和谱范数向外漂移得更快(t^{1/2} 与 t^{1/4})。进一步观察到谱范数扰动有一个非负二阶项。这意味着一个轻量级的“谱帽”——从每次更新中仅投影出单个最大奇异方向的一阶增长——可以在不冻结训练的情况下控制输出协方差W K_X W^T:权重通过非最大方向、最大方向旋转和最大方向切换持续学习。将此帽与奇异值谱的最小熵(H-infinity)相关联。然后研究了用μ子训练的三个系统:一个nanoGPT前馈投影、一个64专家的专家混合路由器以及一个bf16 FlashAttention块的查询/键投影。在每种情况下,谱帽都增加了各向同性,并且在边缘情况下——路由器坍缩为单个专家以及一个注意力头几乎发散——防止了具体的失败,同时使验证损失基本不变。强调尺度不变性假设很强,这些小规模结果是初步的,欢迎评论。
英文摘要
Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood. This preliminary report proposes a unified framework built on a single idealizing assumption -- exact scale invariance of the loss under weight rescaling, which holds approximately in normalization-heavy networks. Under this assumption, plain SGD carries a built-in 1/||W|| brake on its update size, whereas Muon's matrix-sign step removes that brake, so both the Frobenius and spectral norms drift outward faster (t^{1/2} versus t^{1/4}). We further observe that the spectral-norm perturbation has a non-negative second-order term. This implies that a lightweight "spectral cap" -- which projects out only the first-order growth of the single top singular direction from each update -- can control the output covariance W K_X W^T without freezing training: the weight keeps learning through non-top directions, top-direction rotation, and top switching. We relate this cap to the min-entropy (H-infinity) of the singular-value spectrum. We then study three systems trained with Muon: a nanoGPT feed-forward projection, a 64-expert mixture-of-experts router, and the query/key projections of a bf16 FlashAttention block. In each case the cap increases isotropy and, at the margins -- a router collapsing to a single expert, and the near-divergence of one attention head -- prevents a concrete failure, while leaving validation loss essentially unchanged. We emphasize that the scale-invariance assumption is strong and that these small-scale results are preliminary; comments are welcome.