RailWave:面向专家并行通信的自适应时空调度
RailWave: Adaptive Spatial and Temporal Scheduling for Expert-Parallel Communication
浏览论文内容
中文总结 AI 辅助
针对专家并行MoE模型中的不规则全对全通信瓶颈,提出基于DeepEP的相位自适应通信层RailWave,通过空间与时间流量整形实现高达5.84倍加速。
中文摘要 AI 辅助
不规则的全对全通信是专家并行混合专家(MoE)模型中的主要瓶颈。即使采用固定的专家路由和放置策略,并行网络通道(Rails)的不均匀利用以及多对一拥塞(incast)也可能限制通信性能。我们提出了RailWave,一种基于DeepEP构建的相位自适应通信层,通过空间和时间流量整形在路由层之下解决这些瓶颈。RailBalance利用源本地信息将源流量重新分配到符合条件的通道上,而一个可复用的、基于拓扑推导的置换调度限制了每个接收端的并发发送者数量,无需为每个通信阶段重建依赖需求的调度。一个轻量级的校准选择器根据每个阶段的流量特征和离线剖析结果选择执行路径。在来自106B GLM-4.5-Air模型的训练派生通信工作负载上,RailWave在H800上相比Native实现了高达5.84倍的加速,在H20上实现了4.36倍的加速。代码可在https://this URL获取。
英文摘要
Irregular All-to-All communication is a major bottleneck in expert-parallel Mixture-of-Experts (MoE) models. Even with fixed expert routing and placement, uneven utilization of parallel network Rails and incast can limit communication performance. We present RailWave, a phase-adaptive communication layer built on DeepEP that addresses these bottlenecks below the routing layer through spatial and temporal traffic shaping. RailBalance redistributes source traffic across eligible Rails using source-local information, while a reusable, topology-derived permutation schedule limits concurrent senders per receiver without rebuilding demand-dependent schedules for each communication phase. A lightweight calibrated selector chooses an execution path according to each phase's traffic characteristics and offline profiling results. On training-derived communication workloads from the 106B GLM-4.5-Air model, RailWave delivers up to 5.84x speedup on H800 and 4.36x on H20 over Native. Code is available at https://github.com/CyberSecurityErial/RailWave-EP.
发表机构
- Sun Yat-sen University(中山大学)
- Southeast University(东南大学)
- Monash University(蒙纳士大学)
- National University of Singapore(新加坡国立大学)
- Shenzhen University of Advanced Technology(深圳先进技术研究院)
- Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。