arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14385cs.LGcs.AI

DeaMoE:用于快速小批量解码的高效MoE结构

DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding

Zewen Jin, Shen Fu, Zeping Duan, Shannon Wang, Weihao Wu, Chengjie Tang, Congkun Ai, Ping Gong, Zijian Dai, Youhui Bai, Cheng Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对小批量解码下MoE模型专家权重加载的瓶颈,提出DeaMoE架构,通过专家分组共享参数与两阶段路由策略提升效率,在多款模型及显卡上实现显著速度提升。

中文摘要 AI 辅助

混合专家(Mixture-of-Experts,MoE)模型已广泛应用于编程助手、实时音视频交互系统等实时交互式应用中。为满足这些场景极低的响应延迟要求,从业者通常采用小批量解码,在此模式下MoE推理受内存限制,且严重受限于专家权重加载。然而,这一瓶颈未得到足够关注,现有解决方案如训练后权重压缩或预训练期间的细粒度专家设计,要么会降低模型准确率,要么会引入额外计算与通信开销。为解决该问题,本文提出DeaMoE,一种解码高效的MoE架构,其中专家被分组为若干部门,同一部门的专家因来自同一专业领域而共享大部分参数,此外每个专家包含少量私有参数以体现其独特性。本文还为DeaMoE设计了定制的两阶段路由策略以避免冗余加载,使DeaMoE在大语言模型(LLM)解码期间大幅提升效率。与普通MoE相比,DeaMoE在A40显卡上的预训练7B模型上,每步加载权重减少最多50.9%,端到端TPOT速度提升最高达1.33倍;在微基准测试中,DeepSeek-V3在A40和H100显卡上的峰值速度提升分别达2.00倍和1.97倍。

英文摘要

Mixture-of-Experts (MoE) models have been widely adopted in real-time interactive applications such as coding assistants, real-time audio-video interaction systems. To meet the extremely low response latency requirements of these scenarios, practitioners commonly employ small-batch decoding, under which MoE inference becomes memory-bound and is severely bottlenecked by expert weight loading. However, this bottleneck has received limited attention, and existing solutions such as post-training weight compression or fine-grained expert design during pre-training either degrade model accuracy or introduce additional computation and communication overhead. To tackle this issue, we propose DeaMoE, a decoding-efficient MoE architecture, in which the experts are grouped into several departments, and the experts belonging to the same department share most parameters since they come from the same professional field, and additionally each expert contains a few private parameters to reflect its uniqueness. Moreover, we design customized two-stage routing strategy for DeaMoE to avoid redundant loading, under which DeaMoE greatly improves the efficiency during LLM decoding. Compared with vanilla MoE, DeaMoE reduces per-step loaded weights by up to 50.9% and achieves up to 1.33 end-to-end TPOT speedup for the pre-trained 7B model on A40, and up to 2.00x and 1.97x peak speedup for DeepSeek-V3 on A40 and H100 in microbenchmarks.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
  • Shanxi University(山西大学)

机构由 AI 辅助整理,请以论文原文为准。

↑