MAPLE:MoE自适应即插即用分层专家分配
MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation
浏览论文内容
中文总结 AI 辅助
MAPLE是一种即插即用框架,可在不修改预训练MoE LLM权重或重新训练的情况下,异质性分配专家预算,在75%预算下超越基线,降低延迟并提升吞吐量,同时提升多项任务准确率。
中文摘要 AI 辅助
稀疏激活的混合专家(Mixture-of-Experts,MoE)Transformer在所有层中统一固定路由专家的数量,这一惯例忽略了已被充分证实的分层冗余异质性。本文证明这种统一性系统地处于次优状态,并提出MAPLE——一种即插即用框架,可在不修改权重或无需重新训练的情况下,在任何预训练MoE大语言模型(LLM)的各层中异质性地重新分配路由专家预算。核心贡献是一种闭式灵敏度引导分配:探究各层对专家数量变化的响应,用三种指标量化灵敏度,并推导解析最优预算分配,将容量导向敏感层,吸收冗余层的减少。该闭式解还通过灵敏度约束遗传搜索优化,以分层灵敏度为先验引导探索,实现更快收敛和更优分配质量。在四个涵盖不同规模和架构的MoE模型上,MAPLE在75%路由专家预算下优于均匀分配和基于剪枝的基线。值得注意的是,在DeepSeek-MoE-16B上,MAPLE仅使用75%的专家,却在ARC-E、ARC-C和BoolQ上超越原始100%专家均匀基线,准确率分别从65.09提升至71.40、48.49提升至51.50、80.03提升至82.38。这些准确率提升转化为部署效率:在SGLang中实现MAPLE,使单GPU端到端服务延迟降低32.2%,吞吐量提升47.4%。结果表明,精心设计的异质性分配比单纯激活更多专家更有效,为提升MoE效率确立了一个有原则且实用的方向。
英文摘要
Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention that ignores the well-documented heterogeneity in layer-wise redundancy. We demonstrate that this uniformity is systematically suboptimal and propose MAPLE, a plug-and-play framework that reallocates the routed-expert budget heterogeneously across layers of any pretrained MoE LLM, without modifying weights or requiring retraining. Our core contribution is a closed-form sensitivity-guided allocation: we probe each layer's response to variation in expert count, quantify sensitivity using three measures, and derive an analytically optimal budget assignment that directs capacity towards sensitive layers and absorbs reductions in redundant layers. This closed-form solution is further refined by a sensitivity-constrained genetic search that uses layer-wise sensitivity as a prior to guide exploration, yielding faster convergence and superior allocation quality. On four MoE models spanning different scales and architectures, MAPLE outperforms uniform and pruning-based baselines under a 75% routed-expert budget. Notably, on DeepSeek-MoE-16B, MAPLE uses only 75% of the experts yet surpasses the original 100% expert-uniform baseline on ARC-E, ARC-C, and BoolQ, improving accuracy from 65.09 to 71.40, 48.49 to 51.50, and 80.03 to 82.38, respectively. These accuracy gains translate into measured deployment efficiency: implementing MAPLE in SGLang reduces single-GPU end-to-end serving latency by 32.2% and improves throughput by 47.4%. These results show that well-designed heterogeneous allocation can be more effective than simply activating more experts, establishing it as a principled and practical axis for improving MoE efficiency.
发表机构
- University of Bristol(布里斯托尔大学)
机构由 AI 辅助整理,请以论文原文为准。