arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Zepp:在宽松平衡约束下加速分布式混合专家(MoE)服务

Zepp: Accelerating Distributed MoE Serving under Relaxed Balance Constraints

Chang Chen, Andrew Yang, Tiancheng Chen, Jiangfei Duan, Xinwei Qiang, Zhongkai Yu, Xiang Fang, Yufei Ding

arXiv 2610.11158首次发表:更新:

发表机构

UCSD; Stanford University; ETH Zürich; NVIDIA(加州大学圣地亚哥分校; 斯坦福大学; 苏黎世联邦理工学院; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Zepp将平衡视为约束而非目标,通过分阶段优化通信等技术,实现分布式MoE服务的高效执行,在7种基准系统中最高获6.68倍MoE层加速、几何平均1.86倍加速。

AI 中文摘要

随着混合专家(MoE)模型规模不断扩大,其服务部署日益依赖于在越来越多设备上的专家并行(EP)技术。然而,倾斜的专家工作负载会导致计算、通信和内存资源出现不平衡,使得负载均衡成为分布式MoE服务中的核心优化目标。我们发现,实现平衡并非无代价:为平衡某一维度而引入的操作本身可能成本高昂或导致其他维度失衡。这促使我们重新思考,将平衡视为一种约束而非优化目标。我们提出Zepp,该方法在对物理资源(即GPU和NIC)施加简化平衡约束的前提下,直接优化分布式MoE服务中的瓶颈通信。Zepp在部署、路由和执行三个阶段逐步优化节点间通信:首先在GPU约束下部署专家副本以减少令牌通信;接着在NIC约束下通过拆分与合并原语重塑通信流;最后对专家计算进行分区和调度,以重叠生成的通信过程。为适应动态工作负载,Zepp在每次迭代中协同协调计算、令牌通信及专家权重迁移。这些设计共同使Zepp能够追求最高效的执行,而非单一维度的平衡。我们实现了Zepp,并与7种最先进的MoE服务系统进行评估对比,实现了最高6.68倍的MoE层加速,且相较于最快的竞争基准,几何平均加速达1.86倍。

英文摘要

As Mixture-of-Experts (MoE) models continue to scale, serving them increasingly relies on expert parallelism (EP) across a growing number of devices. Yet skewed expert workloads create imbalance across computation, communication, and memory, making load balancing a central optimization objective in distributed MoE serving. We observe that balance is not free: operations introduced to balance one dimension can themselves be expensive or imbalanced. This motivates us to rethink balance as a constraint rather than an optimization objective. We present Zepp, which directly optimizes the bottleneck communication in distributed MoE serving subject to simplified balance constraints on physical resources, i.e., GPUs and NICs. Zepp progressively optimizes inter-node communication across placement, routing, and execution. It first places expert replicas to reduce token communication under GPU constraints, then reshapes communication flows through split and merge primitives under NIC constraints, and finally partitions and schedules ex- pert computation to overlap the resulting communication. To adapt to dynamic workloads, Zepp jointly coordinates computation, token communication, and expert-weight movement at each iteration. Together, these designs allow Zepp to pursue the most efficient execution rather than a single-dimension balanced one. We implement Zepp and evaluate it against 7 state-of-the-art MoE serving systems, achieving up to 6.68$\times$ MoE layer speedup and a geometric mean speedup of 1.86$\times$ over the fastest competing baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑