发表机构
University of Southern California; Meta(南加州大学; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大型扩散Transformer的Muon优化器,提出周期性行级Muon,在保留其生成质量优势的同时,大幅降低训练的计算、通信开销与时间。
AI 中文摘要
感知矩阵的优化器Muon通过在奇异方向间平衡更新来改进大模型训练,但其在大型扩散Transformer(DiTs)上的缩放行为与端到端效率仍不明确。我们首先在参数规模从13亿到150亿的DiTs上建立了Muon的缩放行为,结果显示其相对于AdamW的优化性能与生成质量优势在各模型规模下均保持。然而,在大规模场景下,每步优化均执行的5步牛顿-舒尔茨迭代(NS5),结合全动量物化,会引入大量计算与通信开销,抵消Muon的单步效率优势。我们提出了「周期性行级Muon」,该优化器每K步执行一次完整的NS5谱更新,其余步骤则基于当前动量执行低计算与通信成本的行级约束更新。我们还协同设计了分布式实现方案,该方案在非刷新步骤直接对分片动量进行操作,并通过分桶全收集与通信-计算重叠加速谱刷新。在所有规模下,Muon将观测到的最佳生成质量较AdamW提升12.9%至19.1%;与普通Muon相比,周期性行级Muon在13亿至40亿参数模型上的最佳生成质量差距在0.5%以内,在90亿参数模型上则提升4.5%;它将优化器时间减少46.9%至54.3%,端到端单步时间减少15.7%至24.3%,逻辑通信量减少66.7%,同时达到各自最佳生成质量所需的活跃训练时间减少33.7%至64.8%。这些结果表明,周期性行级Muon在保留Muon生成质量优势的同时,为大型DiTs带来了端到端的训练效率提升。
英文摘要
The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion Transformers (DiTs) remain unclear. We first establish Muon's scaling behavior on DiTs from 1.3B to 15B parameters, showing that its optimization and generative quality advantages over AdamW persist across model scales. However, at scale, the 5-step Newton--Schulz iteration (NS5) performed at every optimization step, together with full-momentum materialization, introduces substantial computation and communication overhead that can offset Muon's step-efficiency advantage. We introduce \emph{Periodic Row-wise Muon}, which performs a full NS5 spectral update once every \(K\) steps and applies a low compute and communication cost row-wise constrained update based on the current momentum at the remaining steps. We further co-design a distributed implementation that operates directly on sharded momentum during non-refresh steps and accelerates spectral refreshes through bucketed all-gather and communication--computation overlap. Across all scales, Muon improves the best observed generative quality over AdamW by 12.9--19.1\%. Compared with vanilla Muon, Periodic Row-wise Muon remains within 0.5\% in best generative quality on the 1.3B--4B models and improves it by 4.5\% at 9B. It reduces optimizer time by 46.9--54.3\%, end-to-end step time by 15.7--24.3\%, and logical communication volume by 66.7\%, while reaching its respective best generative quality with 33.7--64.8\% less active training time. These results show that Periodic Row-wise Muon preserves Muon's generative quality advantage while translating it into end-to-end training efficiency for large DiTs.