LLaDA MoE v2:扩展混合专家扩散语言模型
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
该研究明确MoE扩散语言模型的扩展规律,训练30B-A3B的LLaDA MoE v2,其预训练令牌量为Qwen3的65%,经微调后在多数推理编码基准上优于SDAR Chat且接近Qwen3。
中文摘要 AI 辅助
扩散语言模型(dLLMs)为自回归(AR)语言建模提供了另一种选择,但混合专家(MoE)dLLMs的扩展行为仍鲜为人知。我们系统地表征了MoE dLLMs如何随优化超参数、计算分配和架构扩展,确定了与之前报道的AR模型扩展趋势的定量差异。具体而言,在优化方面,最优名义批大小随计算增长更快,而最优学习率随计算衰减更快;在模型-数据分配方面,IsoFLOP分析显示存在轻微的数据侧倾斜:最优令牌预算增长快于激活的模型侧计算;在MoE架构方面,更大规模在固定激活容量下更倾向于更大的专家池,而中等专家粒度始终有效,且分配给共享专家的激活容量的首选比例在各规模间保持稳定。基于这些发现,我们从头开始在23.5T令牌上训练了30B-A3B的dLLM LLaDA MoE v2。使用约为Qwen3 65%的预训练令牌,LLaDA MoE v2在多个知识、推理和编码基准上接近Qwen3;仅经过监督微调后,它在8个推理和编码基准中的7个上优于SDAR Chat,且在多个任务上仍接近Qwen3。这些结果为MoE dLLMs确立了实用的扩展定律和设计原则。
英文摘要
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Specifically, for optimization, the optimal nominal batch size grows faster, while the optimal learning rate decays more rapidly with compute. For model--data allocation, IsoFLOP analysis reveals a slight data-side tilt: the optimal token budget grows faster than activated model-side computation. For MoE architecture, larger scales increasingly favor larger expert pools at fixed activated capacity, while moderate expert granularity remains consistently effective and the preferred fraction of activated capacity assigned to shared experts remains stable across scales. Guided by these findings, we train LLaDA MoE v2, a 30B-A3B dLLM, from scratch on 23.5T tokens. With approximately 65\% as many pretraining tokens as Qwen3, LLaDA MoE v2 approaches Qwen3 on several knowledge, reasoning, and coding benchmarks. After supervised fine-tuning alone, it outperforms SDAR Chat on seven of eight reasoning and coding benchmarks and remains close to Qwen3 on several tasks. These results establish practical scaling laws and design principles for MoE dLLMs.