UMoE:在特定领域训练中解锁每个专家
UMoE:Unlocking Every Expert in Domain-Specific Training
浏览论文内容
中文总结 AI 辅助
研究针对特定领域训练中专家池不合理问题,提出UMoE流程,先修剪低显著性专家,再扩展专家池,最后应用标准SFT,该方法在多架构、多领域及多基准测试中优于直接SFT,能转化冗余容量,提升训练效果。
中文摘要 AI 辅助
混合专家(MoE)模型在不按比例增加计算成本的情况下扩展容量,已成为前沿大语言模型的关键架构。然而,特定领域的训练后继承了由混合领域预训练塑造的专家池:相当一部分专家在目标领域贡献很小,标准监督微调(SFT)使该专家池的组成保持不变。我们提出了一种简单的、保持预算的流程,在微调前将专家池重新调整到目标领域。给定一个目标领域,我们(1)修剪领域对齐显著性最低的专家,(2)通过基于扰动的专家扩展将专家池重新增长到原来的大小,(3)应用标准SFT。得到的模型保留了原来的专家数量、参数数量和推理成本。UMoE通过单一的冻结方法且无需每个领域的超参数调整,在两种MoE架构(Qwen3 - 30B - A3B和Qwen3.5 - 35B - A3B)、五个领域(数学、代码、科学、工具使用和智能编码)以及12个基准测试中始终优于直接SFT。在强大的内部数学语料库上,直接SFT已经超过Qwen3 - 30B - A3B - Thinking(82.81对81.06),而UMoE进一步将平均值提高到84.17,额外提高了1.36分,证明了对更强SFT机制的鲁棒性。数据缩放实验进一步表明,随着训练数据的增加,增益仍然存在。分析表明,直接SFT模型将大量路由专家计算分配给一个低显著性子集,该子集可以事后去除而平均降级很小;UMoE将这种冗余容量转化为有用的领域容量,并实现了更低的训练损失,在下游评估中所有难度级别都有增益。
英文摘要
Mixture-of-Experts (MoE) models scale capacity without proportional compute cost and have become a key architecture for frontier large language models (LLMs). Yet domain-specific post-training inherits an expert pool shaped by mixed-domain pre-training: a substantial subset of experts contributes little on the target domain, and standard supervised fine-tuning (SFT) leaves the composition of this pool unchanged. We propose a simple, budget-preserving pipeline that realigns the expert pool to the target domain before fine-tuning. Given a target domain, we (1) prune the experts with lowest domain-aligned saliency, (2) regrow the expert pool to its original size through perturbation-based expert expansion, and (3) apply standard SFT. The resulting model preserves the original expert count, parameter count, and inference cost. With a single frozen recipe and no per-domain hyperparameter tuning, UMoE consistently improves over direct sft across two MoE architectures (Qwen3-30B-A3B and Qwen3.5-35B-A3B), five domains (math, code, science, tool-use, and agentic coding), and 12 benchmarks. Representative improvements are 3.4 points in math average accuracy, 6.0 points on SWE-bench Verified. On a strong in-house math corpus, direct sft already surpasses Qwen3-30B-A3B-Thinking (82.81 vs.\ 81.06), yet UMoE further raises the average to 84.17, an additional 1.36 points, demonstrating robustness to a substantially stronger SFT regime. Data-scaling experiments further show that the gain persists as training data grows. Analysis reveals that the direct-SFT model allocates substantial routed-expert compute to a low-saliency subset that can be removed post hoc with little average degradation; UMoE turns this redundant capacity into useful domain capacity and achieves lower training loss, with gains spanning all difficulty levels in downstream evaluation.
发表机构
- Qwen Team(通义千问团队)
- DeepSeek-AI(渊亭科技)
- Thinking Machines Lab(思维机器实验室)
- LLM-Core Xiaomi(小米大语言模型核心团队)
机构由 AI 辅助整理,请以论文原文为准。