Purlin:将编排与集合通信的数据路径分离
Purlin: Separating Orchestration from the Datapath of Collectives
浏览论文内容
中文总结 AI 辅助
Purlin通过分离编排与数据路径,实现了GPU集体通信的定制化与硬件适配,在多种GPU上显著提升性能,并集成到SGLang中改善LLM服务与图像生成效率。
中文摘要 AI 辅助
分布式推理依赖于GPU集体通信,这种通信必须跟上不断发展的硬件和专门工作负载。然而,现有的集体通信实现常常将语义、编排(数据在何处以及何时移动)和数据路径(数据如何移动)耦合在一起。这种耦合使得采用新的硬件机制和为应用定制通信变得代价高昂。我们提出了Purlin,一个规模扩展通信框架,将这些关注点分离开来。在Purlin的顶层,我们将集体通信指定为输入和输出布局的命名以及复制或归约操作。在中间层,我们引入了一个共享的编排协议,即阶段、通知和消费(SNAC),它从这些规范中推导出协调。在SNAC之下是一个特定于硬件的数据路径,我们称之为Atom,它实现了集体通信的两个关键数据移动原语:复制和归约。这种分离使我们能够定制集体通信并采用新的硬件机制,同时通过SNAC重用编排。我们在A100、H200和B200 GPU上评估了Purlin。在七种集体通信中,Purlin相对于基线实现了高达5.14倍的延迟加速和高达4.50倍的带宽提升。集成到SGLang中,Purlin将离线LLM服务吞吐量和交互性平均提高了1.13倍,最高可达1.37倍。对于在线LLM推理,Purlin将交互性平均提高了1.26倍,最高可达2.85倍,最大的增益出现在过载情况下。对于扩散图像生成,Purlin将端到端延迟降低了高达1.13倍。
英文摘要
Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.
发表机构
- Stanford University(斯坦福大学)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。