arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Director:通过在线主动专家放置加速分布式混合专家服务

Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement

Qianli Liu, Kaibin Guo, Zicong Hong, Peng Li, Fahao Chen, Haodong Wang, Jian Lin, Song Guo

arXiv 2607.08782首次发表:更新:

AI 中文总结

研究针对混合专家模型服务中专家放置问题,提出Director系统,通过预测驱动的在线专家放置、轻量级预测器或量化副本及在线迁移模块,结合多项式时间运行的优化器,有效降低端到端延迟,实验证明其对流行MoE模型延迟降低显著。

AI 中文摘要

专家并行已成为服务混合专家(MoE)模型的主流范式,其效率取决于GPU的通信和计算延迟,而这与专家在GPU中的放置有关。现有优化专家放置的工作侧重于利用过去请求的专家激活模式,但面对多样化和快速变化的请求模式存在不足。我们提出了Director,一种新的分布式MoE服务系统,通过预测驱动的在线专家放置最小化端到端延迟。Director使用轻量级级联预测器或低位量化副本预测传入请求的专家激活模式,在线迁移模块在计算受限阶段执行迁移以实现近零停机时间。核心的基于松弛的专家放置优化器在容量约束下运行,多项式时间内实现(1 + ε)近似比率。通过广泛实验表明,与现有工作相比,流行MoE模型的端到端延迟降低了11%至55%。

英文摘要

Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs. Existing works for optimizing expert placement focus on leveraging past requests' expert activation patterns. However, they demonstrate deficiencies facing diverse and rapidly changing request patterns, calling for an online, proactive approach. Implementing such an approach requires addressing several challenges: the uncertainty associated with incoming requests' expert activation, the cost of expert migration, and the NP-hard complexity in optimization. Therefore, we present Director, a new distributed MoE serving system that minimizes end-to-end latency via prediction-driven, online expert placement. Director uses either a lightweight cascaded predictor or a low-bit quantized replica for expert activation patterns of incoming requests. An online migration module then enacts the changes with near-zero downtime by executing migrations in compute-bound phases, keeping disruption bounded. At its core, a relaxation-based expert placement optimizer operates under capacity constraints, runs in polynomial time, and achieves a $(1+ε)$ approximation ratio. Finally, we implement a prototype and demonstrate, through extensive experiments, a reduction in end-to-end latency of $11\sim55\%$ for popular MoE models (e.g., Mistral, DeepSeek and Qwen) compared to existing work.

CommentsINFOCOM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑