VarioPath:面向PCIe GPU集群的工作负载感知的全对全通信
VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters
浏览论文内容
中文总结 AI 辅助
VarioPath提出结合离线拓扑分析与在线需求感知的调度框架,解决PCIe GPU集群中AlltoAllv通信的链路争用问题,实现平均5.88倍加速并显著降低推理延迟。
中文摘要 AI 辅助
AlltoAllv通信是分布式大模型推理中的关键原语,尤其在混合专家(MoE)模型中。随着PCIe GPU系统在成本效益推理中的日益普及,AlltoAllv在这些系统上的性能变得愈发重要。在没有专用横向扩展互连(如NVLink或Infinity Fabric)的情况下,PCIe GPU系统通过PCIe层次结构承载节点内和节点间流量,并发传输可能争用PCIe链路带宽。这种链路争用,加上偏斜的流量分布和动态流量需求,使得高效的AlltoAllv调度变得具有挑战性。现有方法要么不适合PCIe GPU系统,要么产生大量的调度综合开销,降低了其在实际部署中的实用性。我们提出了VarioPath,一种针对PCIe GPU系统的高效AlltoAllv调度框架。它结合了离线拓扑感知分析器和在线需求感知调度器。分析器将无争用传输模式记录为AlltoAllv通道,并利用拓扑对称性构建紧凑的目录以进行高效搜索。在线调度器将每次AlltoAllv调用的需求分解为一系列通道,适应快速变化和偏斜的流量,同时保持较低的计划开销。在四个平台(最多256个GPU)上的评估显示,与FAST相比,平均AlltoAllv加速比为5.88倍,与DeepEP相比为1.72倍。端到端实验表明,VarioPath将Qwen3推理延迟降低了高达27.2%,将Wan2.1生成延迟降低了6.1%。
英文摘要
AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e.g., NVLink or Infinity Fabric), PCIe GPU systems carry both intra-node and inter-node traffic through the PCIe hierarchy, where concurrent transfers can contend for PCIe link bandwidth. This link contention, compounded by skewed traffic distributions and dynamic traffic demand, makes efficient AlltoAllv scheduling challenging. Existing approaches are either poorly suited to PCIe GPU systems or incur substantial schedule synthesis overhead that reduces their practicality in real-world deployments. We present VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems. It combines an offline topology-aware analyzer with an online demand-aware scheduler. The analyzer records contention-free transfer patterns as AlltoAllv channels and exploits topology symmetry to build a compact catalog for efficient search. The online scheduler decomposes each AlltoAllv invocation's demand across a sequence of channels, adapting to rapidly changing and skewed traffic while incurring low planning overhead. Evaluation on four platforms (up to 256 GPUs) shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP. End-to-end experiments show that VarioPath reduces Qwen3 inference latency by up to 27.2% and Wan2.1 generation latency by 6.1%.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。