发表机构
Peking University; Westlake University; Zhejiang University(北京大学; 西湖大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散Transformer训练计算成本高、普通Muon优化器存在后期收敛问题的瓶颈,提出CMuon策略,通过分块权重实现训练加速与收敛优化,在ImageNet 256上200个epoch达FID1.18,训练速度超AdamW2倍
AI 中文摘要
扩散Transformer(DiTs)在视觉生成建模中已实现了当前最优(SOTA)性能,但其训练仍存在计算成本过高的问题。尽管近期提出的动量正交化(Muon)优化器为AdamW提供了颇具前景的替代方案,但将其直接应用于DiTs会导致后期收敛效果欠佳。在本文中,我们确定了这一瓶颈的根本原因:标准DiT架构为提升计算效率,将功能不同的权重(例如AdaLN和QKV层内的权重)融合为统一张量。将Muon应用于这些融合张量会无意间引发隐式子空间耦合,从而扭曲更新方向并降低全局优化效果。为解决这一问题,我们提出了分块Muon(CMuon),这是一种简单却极为有效的策略,即在正交化之前将这些矩阵划分为独立的子组件。大量实验表明,采用CMuon训练的6.75亿参数DiT在ImageNet 256数据集上仅用200个epoch就达到了1.18的FID值。这相比AdamW实现了超过2倍的训练加速,同时有效克服了普通Muon的后期收敛平台问题。
英文摘要
Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence. In this paper, we identify the root cause of this bottleneck: standard DiT architectures fuse functionally distinct weights (e.g., within AdaLN and QKV layers) into unified tensors for computational efficiency. Applying Muon to these fused tensors inadvertently induces implicit subspace coupling, which distorts update directions and degrades global optimization. To address this, we introduce Chunked Muon (CMuon), a simple yet highly effective strategy that partitions these matrices into independent sub-components prior to orthogonalization. Extensive experiments demonstrate that a 675M-parameter DiT trained with CMuon achieves a FID of 1.18 on ImageNet 256 in just 200 epochs. This represents more than a 2x training speedup over AdamW, while effectively overcoming the late-stage convergence plateaus of vanilla Muon.
CommentsECCV 2026