GPU服务器拓扑、并行度与拥塞控制对MoE推理的联合影响:一项受控仿真研究
Joint Effects of GPU Server Topology, Parallelism, and Congestion Control on MoE Inference: A Controlled Simulation Study
查看机构详情
- Inspur Group(浪潮集团)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过768次确定性仿真,系统探究了GPU服务器拓扑、并行度及拥塞控制对MoE推理完成时间的联合影响,发现这些因素共同决定暴露通信与性能,并验证了配置满足权重驻留约束。
中文摘要 AI 辅助
混合专家(MoE)模型通过稀疏激活扩展容量,但跨GPU推理引入了张量并行(TP)集合通信以及专家并行(EP)的分发和合并操作。完成时间不仅取决于通信量,还取决于逻辑组如何映射到服务器内部互连、GPU-NIC连接以及节点间网络。使用带有NS-3离散事件后端的ASTRA-sim,我们构建了一个包含32个GPU秩的受控矩阵,其中数据和流水线并行度固定为一。工作负载是来自四种MoE配置的固定长度4096令牌的类预填充合成Chakra轨迹。实验覆盖六种服务器拓扑、四种TP/EP分区、两种TP集合通信算法以及四种网络/拥塞控制模式,共产生768次确定性仿真。在每个模型的144配置反馈启用子集中,暴露的通信占平均完成时间的89.9%至95.8%。TP16EP2所需的平均完成时间是TP2EP16的3.68至4.35倍。在固定秩映射下,ASTRA-sim双二叉树(DBT)比环(Ring)多花费28.3%至83.2%的时间。类似InfiniBand的高精度拥塞控制(HPCC)比基于融合以太网的RDMA(RoCE)上的HPCC低约0.9%,而带有数据中心量化拥塞通知(DCQCN)的RoCE比RoCE HPCC慢23.8%至35.7%。拓扑效应是有条件的:拓扑6在低TP度下领先,但在高TP度下失去优势,并且额外的GPU或NIC仅在秩映射平衡注入路径上的流量时才有帮助。在统一的32路分片下,最大的检查点权重分片约为每秩48.75 GB,因此所有配置都满足每加速器64 GB的权重驻留标准。在评估的工作负载和仿真器语义内,服务器拓扑、并行度、集合通信实现和拥塞控制共同决定暴露的通信和完成时间。
英文摘要
Mixture-of-experts (MoE) models expand capacity via sparse activation, but inference across GPUs introduces tensor-parallel (TP) collectives and expert-parallel (EP) dispatch and combine operations. Completion time depends not just on communication volume but on how logical groups map onto intra-server interconnects, GPU--NIC connections, and the inter-node network. Using ASTRA-sim with the NS-3 discrete-event backend, we build a controlled matrix of 32 GPU ranks with data and pipeline parallelism fixed at one. Workloads are fixed-length 4096-token prefill-like synthetic Chakra traces from four MoE configurations. Experiments cover six server topologies, four TP/EP partitions, two TP collective algorithms, and four network/congestion-control modes, yielding 768 deterministic simulations. In the 144-configuration feedback-enabled subset per model, exposed communication accounts for 89.9%--95.8% of mean completion time. TP16EP2 requires 3.68--4.35x the mean completion time of TP2EP16. With fixed rank mapping, ASTRA-sim Double Binary Tree (DBT) incurs 28.3%--83.2% more time than Ring. InfiniBand-like High Precision Congestion Control (HPCC) is ~0.9% lower than HPCC over RDMA over Converged Ethernet (RoCE), whereas RoCE with Data Center Quantized Congestion Notification (DCQCN) is 23.8%--35.7% slower than RoCE HPCC. Topology effects are conditional: Topology~6 leads at low TP degrees but loses its advantage at high TP degrees, and additional GPUs or NICs help only when rank mapping balances traffic across injection paths. Under uniform 32-way sharding, the largest checkpoint-weight shard is ~48.75 GB per rank, so all configurations meet a 64 GB per-accelerator weight-residency criterion. Within the evaluated workload and simulator semantics, server topology, parallelism, collective implementation, and congestion control jointly determine exposed communication and completion time.