arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DivMoE:通过跨领域专家组合实现细粒度MoE升级

DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition

Yuxuan Lou, Kai Yang, Geng Zhang, Yong Liu, Yang You

arXiv 2610.11317首次发表:更新:

发表机构

School of Computing, National University of Singapore; Shanghai Jiao Tong University(新加坡国立大学计算机学院; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DivMoE是首个实现细粒度MoE升级的框架,通过领域专用专家初始化与多样性约束路由,在15个基准测试中优于6个升级基线,性能接近更大规模的Moonlight-MoE模型。

AI 中文摘要

混合专家(MoE)架构已成为扩展大语言模型的核心,近期研究表明细粒度专家设计具有显著优势。从头训练此类模型成本高昂,而从预训练稠密模型进行稀疏升级是颇具吸引力的替代方案。但我们发现细粒度升级存在结构缺陷:当细粒度专家源自单一源模型时,简单路由会发生崩溃,下游准确率降至接近随机水平(例如在Qwen3-1.7B上,Drop-Upcycling-fine-grained在15个基准测试中平均准确率仅为23.2%,与从头训练的22.2%基本相当,而该方法的粗粒度变体准确率达50.2%)。我们提出DivMoE,这是首个实现结构平衡路由的细粒度MoE升级框架。DivMoE引入领域专用细粒度专家初始化,从经领域自适应持续预训练的稠密模型中派生专家;还引入多样性约束路由,这是一种硬性结构约束,保证每个token激活来自不同领域组的专家。在两个基础模型和15个基准测试中,DivMoE始终优于6个升级基线(在Qwen3-1.7B上,与最强基线相比为55.6% vs. 51.6%),且在第二阶段持续预训练后,在每个基准测试中均严格优于稠密基础模型——消除了困扰先前细粒度升级的性能退化差距。在公开推理混合数据集上进行监督微调后,我们的120亿参数DivMoE模型在平均准确率64.5%上与Moonlight-MoE(160亿参数)相当,同时比受控NVIDIA-Upcycling基线高出6.1个百分点。

英文摘要

Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained expert designs. Training such models from scratch is expensive, and sparse upcycling from pre-trained dense models is an attractive alternative. However, we identify a structural pathology of fine-grained upcycling: when fine-grained experts are derived from a single source model, naive routing collapses and downstream accuracy drops to near-random (e.g., on Qwen3-1.7B, Drop-Upcycling-fine-grained reaches only 23.2% average accuracy across 15 benchmarks, essentially matching from-scratch training at 22.2%, while the same method's coarse-grained variant reaches 50.2%). We propose DivMoE, the first framework achieving fine-grained MoE Upcycling with structurally-balanced routing. DivMoE introduces domain-specialized fine-grained expert initialization, deriving experts from dense models that have undergone domain-adaptive continual pre-training, and diversity-constrained routing, a hard structural constraint guaranteeing that each token activates experts from distinct domain groups. Across two base models and 15 benchmarks, DivMoE consistently outperforms six upcycling baselines (55.6% vs. 51.6% for the strongest baseline on Qwen3-1.7B) and strictly improves over the dense base model on every benchmark after Stage 2 continual pre-training -- closing the regression gap that has plagued prior fine-grained upcycling. After supervised fine-tuning on a public reasoning mixture, our 12B-parameter DivMoE model matches Moonlight-MoE (16B) at 64.5% average accuracy while outperforming a controlled NVIDIA-Upcycling baseline by 6.1 percentage points.

Comments19 pages, published as a conference paper at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑