arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MMOE:通过高效专家设计使扩散变压器现代化

MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang, Erik Cambria, Xuelong Li

arXiv 2607.24665首次发表:更新:

发表机构

College of Computing and Data Science, Nanyang Technological University; Institute of Artificial Intelligence of China Telecom (TeleAI)(南洋理工大学计算与数据科学学院; 中国电信人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究AIGC基础模型中扩散变压器未平衡质量与成本问题,提出MMOE对其现代化改造,将多种专家设计应用于AIGC生成,实验表明MMOE能达更好质量成本平衡,使AIGC基础模型可循大语言模型平衡扩展路径。

AI 中文摘要

现代大语言模型通过将容量增长与效率相结合来成功扩展,控制每令牌和部署成本。AIGC基础模型,尤其是扩散变压器主干,已开始采用稀疏专家,但近期工作大多增加了总参数数量和稀疏率,却未引入使大语言模型扩展实用的效率机制,导致生成质量与训练和部署成本失衡。本文提出疑问:高效大语言模型扩展背后的架构原则能否以更平衡的方式应用于AIGC基础模型?为此引入ModernMOE(MMOE),对SiT风格的扩散变压器进行现代化改造,系统地将路由专家、共享和轻量级专家、门残差路由以及注意力残差信息重用应用于AIGC生成。通过在单个八GPU H100节点上以批量大小256训练400k步的实验,MMOE在每个记录的检查点达到更低的FID,比密集和中间稀疏专家基线收敛更快,在稀疏变体中实现了最佳的质量成本平衡。路由分析还显示了跨深度的稳定专家专业化、轻量级路由的大量使用以及去噪过程中适度的逐步骤路由变化。这些结果表明,AIGC基础模型可以通过引入经过验证的效率设计来遵循大语言模型的平衡扩展路径,而不是简单地增加总参数和稀疏率。

英文摘要

Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly enlarge total parameter counts and sparsity ratios without importing the efficiency mechanisms that made LLM scaling practical, so generation quality is seldom balanced against training and deployment cost. This raises a natural question: can the architectural principles behind efficient LLM scaling be adapted to AFMs in a more balanced way? We introduce ModernMOE (MMOE), a modernization of SiT-style diffusion transformers that systematically adapts routed experts, shared and lightweight experts, gate-residual routing, and attention-residual information reuse to AIGC generation. Rather than treating MoE as a single plug-in replacement, MMOE studies how different modern expert components affect convergence, efficiency, and generation quality when composed inside a diffusion transformer. Every experiment in this paper is trained on a single eight-GPU H100 node with batch size 256 for 400k steps, an accessible single-machine budget. Under matched training and sampling protocols and at this budget, MMOE reaches lower FID at every recorded checkpoint, that is, it converges faster per training step, than dense and intermediate sparse-expert baselines, and among the sparse variants it attains the best quality-cost balance. Routing analysis further shows stable expert specialization across depth, substantial use of lightweight routes, and modest step-to-step routing changes during denoising. These results suggest that AFMs can follow the balanced scaling path of LLMs by importing proven efficiency designs, rather than by simply increasing total parameters and sparsity ratios.

Comments13 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑