arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.16981cs.LGmath.OC

MuonBP:通过分块周期正交化实现更快的 Muon

MuonBP: Faster Muon via Block-Periodic Orthogonalization

  • AWS(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

Ahmed Khaled, Kaan Ozkara, Tao Yu, Mingyi Hong, Youngsuk Park

更新

AI总结:

针对模型并行下 Muon 正交化带来的跨设备通信开销,本文提出 MuonBP,通过局部分块正交化与周期性完整正交化及双步长理论保证,在 8B 模型训练中较 Muon 提升 8% 吞吐量且不损失性能。

AI中文摘要:

梯度正交化是一种简单策略,在加速梯度下降方面表现出很强的实用性。Muon 优化器(Jordan、Jin 等人,2024)将梯度正交化与一阶梯动量相结合,在语言模型训练中相较 Adam/AdamW(Loshchilov 和 Hutter,2019)显著提升了数据效率。然而,在使用模型并行时,由于需要对来自不同设备的梯度矩阵分片执行额外的 gather 和 scatter 操作,梯度正交化相较逐坐标优化器(如 AdamW)会引入额外开销。与 Adam/AdamW 相比,这部分额外通信可能导致 5%–10% 的吞吐量损失。为解决该问题,我们提出具有分块周期正交化的 Muon(Muon with Block-Periodic Orthogonalization,MuonBP):它对每个设备上的矩阵分片独立执行正交化,并周期性地执行完整正交化,以在大规模训练中维持稳定性。我们展示了如何将基线学习率调整为适用于 MuonBP 的学习率,并给出该算法的收敛保证。关键在于,我们的理论要求使用两种步长:一种用于分块正交化步骤,另一种用于完整正交化步骤。我们的方法简单,只需极少的超参数调整,并且在迭代复杂度上与基线 Muon 具有竞争力,同时每轮迭代的吞吐量可与 AdamW 等逐坐标方法相当。在使用八路张量并行和 ZeRO 优化器状态分片训练一个 8B 模型时,MuonBP 相较 Muon 实现了 8% 的吞吐量提升,且性能没有下降。

英文摘要:

Gradient orthogonalization is a simple strategy that shows great utility in speeding up gradient descent. The Muon optimizer (Jordan, Jin, et al., 2024) combines gradient orthogonalization with first-order momentum and achieves significant improvement in data efficiency over Adam/AdamW (Loshchilov and Hutter, 2019) for language model training. However, when using model parallelism, gradient orthogonalization introduces additional overhead compared to coordinate-wise optimizers (such as AdamW) due to additional gather and scatter operations on gradient matrix shards from different devices. This additional communication can amount to a throughput hit of 5%-10% compared to Adam/AdamW. To remedy this, we propose Muon with Block-Periodic Orthogonalization (MuonBP), which applies orthogonalization independently to matrix shards on each device and periodically performs full orthogonalization to maintain training stability at scale. We show how to adjust the learning rate from the baseline to MuonBP and give convergence guarantees for this algorithm. Crucially, our theory dictates that we use two stepsizes: one for the blockwise orthogonalization steps, and one for the full orthogonalization steps. Our method is simple, requires minimal hyperparameter adjustments, and achieves competitive iteration complexity compared with baseline Muon while providing per-iteration throughput comparable to coordinate-wise methods such as AdamW. When training an 8B model with eight-way tensor parallelism and ZeRO optimizer state sharding, MuonBP achieves 8% throughput increase compared to Muon with no degradation in performance.

↑