CascadeEP:注意力不平衡下MoE预填充的异步专家执行
CascadeEP: Asynchronous Expert Execution for MoE Prefill under Attention Imbalance
- University of the Chinese Academy of Sciences(中国科学院大学)
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
- Peking University(北京大学)
- Beijing University of Posts and Telecommunications(北京邮电大学)
- University of Leeds(利兹大学)
- Advanced Institute of Information Technology(先进信息技术研究院)
- Huawei Technologies Ltd.(华为技术有限公司)
- Beijing Tongming Lake Information Technology Application Innovation Center(北京通明湖信息技术应用创新中心)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对MoE预填充中注意力不平衡导致的同步EP延迟,提出异步专家执行引擎ASYNCEP,通过异步EP、streamFFN和OEWF机制,实现高达1.48倍TTFT加速和1.17倍吞吐提升。
中文摘要 AI 辅助
混合专家(MoE)服务通常采用数据并行和专家并行(DEP):注意力副本运行不同的请求批次,而路由专家在专家并行(EP)组中进行分片。在预填充阶段,注意力副本在不同时间完成调度,但同步EP会延迟专家前馈网络(FFN)计算,直到所有副本的路由输入就绪。请求调度器寻求在重用共享提示前缀的键值(KV)缓存以平衡负载,同时避免冗余的预填充计算。当持有匹配前缀的副本已经过载时,这些目标可能冲突,留下残余的注意力不平衡。我们提出了ASYNCEP,一个用于MoE预填充的分布式执行引擎。ASYNCEP提出了三种机制。异步EP允许专家计算在来自所有注意力副本的令牌就绪之前开始。streamFFN批量处理就绪的令牌,以平衡早期执行与FFN计算效率。机会性专家权重获取(OEWF)允许较快的副本获取专家权重并执行来自其他GPU的未启动工作。我们在DeepSeek-V4-Flash、DeepSeek-V4-Pro和GLM-5.3上评估了ASYNCEP,结果显示ASYNCEP在p95首令牌时间(TTFT)上实现了高达1.48倍的加速,并将推理吞吐量提高了高达1.17倍。
英文摘要
Mixture-of-experts (MoE) serving commonly deploys data and expert parallelism (DEP): attention replicas run distinct request batches while routed experts are sharded across an expert-parallel (EP) group. During prefill, attention replicas finish dispatch at different times, but synchronous EP delays expert feed-forward network (FFN) computation until routed inputs from all replicas are ready. Request schedulers seek to balance load while reusing the key-value (KV) cache of shared prompt prefixes to avoid redundant prefill computation. These goals can conflict when a replica holding a matching prefix is already overloaded, leaving residual attention imbalance. We present ASYNCEP, a distributed execution engine for MoE prefill. ASYNCEP proposes three mechanisms. Asynchronous EP allows expert computation to start before tokens from all attention replicas are ready. streamFFN batches ready tokens to balance early execution with FFN computation efficiency. Opportunistic expert weight fetching (OEWF) allows a faster replica to fetch expert weights and execute unstarted work from other GPUs. We evaluate ASYNCEP on DeepSeek-V4-Flash, DeepSeek-V4-Pro, and GLM-5.3, and our results show that ASYNCEP achieves up to 1.48x speedup in p95 time-to-first-token (TTFT) and improves the inference throughput by up to 1.17x.