发表机构
KTH Royal Institute of Technology(瑞典皇家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究 Muon 优化器是否需要细粒度频谱整形,提出双频带重加权框架 BulkBoost,证明单个体到尖峰增益即可获得与细粒度轮廓相当的收益,并降低损失。
AI 中文摘要
Muon 将当前梯度和历史梯度组合成矩阵动量。对于 $M=U\Sigma V^\top$,理想化的极坐标更新 $Q=UV^\top$ 赋予每个奇异方向相同的权重。我们称之为平坦轮廓。最近的几种优化器用细粒度频谱映射取代了这一平坦轮廓,为每个方向提供各自的增益。我们探究 Muon 更新需要多少这种频谱细节。我们的频谱诊断显示,约 $94$--$97\\%$ 的测量奇异模式位于估计的噪声边缘之下,但整体上与参考梯度正向对齐。我们引入了 BulkBoost,一种双频带频谱重加权框架,具有固定秩和噪声校准两种变体。后者利用分割小批量梯度差异,为 Muon 的 Nesterov 输入校准 Marchenko--Pastur 参考边缘,将边缘下方的体部分与上方的尖峰部分分开。两种变体通过一个共享增益增加体部分的相对权重,同时保持每个矩阵未重加权方向的 Frobenius 范数。对于固定划分,我们的理论给出了将权重移向体部分可降低损失的一阶条件。它还量化了在所有逐模式重新分配中,两个频带能够捕获的最大一阶改进率比例。在跨越 Pythia-14M 到 410M 及六个语料库的 30 个持续预训练设置中,双频带重加权与 Freon 的细粒度幂律轮廓相当,并优于 Spectra。与 Muon 的平坦轮廓相比,Freon 平均将最终损失降低了预适应损失的 $0.022\\%$,而双频带变体实现了 $0.073$--$0.147\\%$ 的降低。这些观察表明,偏离平坦轮廓的有用变化出乎意料地低维:单个体到尖峰增益至少能捕获与细粒度频谱轮廓一样多的益处。
英文摘要
Muon combines current and past gradients into matrix momentum. For $M=UΣV^\top$, the idealized polar update $Q=UV^\top$ gives every singular direction the same weight. We refer to this as the flat profile. Several recent optimizers replace this flat profile with fine-grained spectral maps that give each direction its own gain. We ask how much of this spectral detail a Muon update needs. Our spectral diagnostics show that approximately $94$--$97\%$ of measured singular modes lie below an estimated noise edge, yet collectively align positively with a reference gradient. We introduce BulkBoost, a two-band spectral reweighting framework with fixed-rank and noise-calibrated variants. The latter uses split-minibatch gradient differences to calibrate a Marchenko--Pastur reference edge for Muon's Nesterov input, separating the bulk below the edge from the spikes above it. Both variants increase the bulk's relative weight through one shared gain while preserving the Frobenius norm of each matrix's unreweighted direction. For a fixed partition, our theory gives the first-order condition under which moving weight toward the bulk lowers the loss. It also quantifies the fraction of the maximal first-order improvement rate, over all per-mode reallocations, that two bands can capture. Across 30 continued-pretraining settings spanning Pythia-14M to 410M and six corpora, two-band reweighting is competitive with the fine-grained power-law profile of Freon and outperforms Spectra. Measured against Muon's flat profile, Freon reduces final loss by $0.022\%$ of the pre-adaptation loss on average, whereas the two-band variants achieve reductions of $0.073$--$0.147\%$. These observations suggest that useful departures from the flat profile are surprisingly low-dimensional: a single bulk-to-spike gain captures at least as much benefit as the fine-grained spectral profiles.