SparSP:利用通信稀疏性实现序列并行视频扩散Transformer
SparSP: Exploiting Communication Sparsity for Sequence-Parallel Video DiTs
查看机构详情
- University of Waterloo(滑铁卢大学)
- Rice University(莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出SparSP,一种利用注意力稀疏性作为通信原语的序列并行系统,通过依赖感知放置、需求导向KV路由和解耦传输运行时,在PCIe多GPU环境下实现视频扩散Transformer的高效推理,提升注意力性能1.38-1.5倍,端到端加速最高1.69倍。
中文摘要 AI 辅助
扩散Transformer已成为视频生成的主导架构。其巨大的计算成本促使在多GPU服务器上进行推理扩展,但在通过PCIe连接的商用GPU上,由于带宽有限,高效扩展仍然具有挑战性。尽管稀疏注意力大幅减少了计算量,但其对通信的影响仍未得到充分探索。本文认为,注意力稀疏性应被视为一种通信原语。我们提出了SparSP,一种高效的稀疏序列并行通信系统,它协同设计了令牌分布、通信路由和异步执行,以支持稀疏视频扩散模型。首先,依赖感知放置根据扩散模型的稀疏注意力模式分配序列块。其次,需求导向的KV路由将KV块直接传输到请求的GPU,无需中间中继。第三,解耦传输运行时将通信与GPU计算分离,以减少资源争用并最大化有效带宽。我们的评估表明,SparSP将注意力性能提升了1.38-1.5倍,在三个代表性服务器和三个视频扩散模型上实现了平均1.17倍(最高1.69倍)的端到端加速,并将通信量减少了12.54-23.05%。此外,与NCCL相比,我们实现了平均1.53-1.76倍的带宽提升。
英文摘要
Diffusion Transformers have become the dominant architecture for video generation. Their substantial computational cost motivates scaling inference across multi-GPU servers, yet efficient scaling remains challenging on commodity GPUs connected via PCIe, whose bandwidth is limited. Although sparse attention substantially reduces computation, its implications for communication remain underexplored. This paper argues that attention sparsity should be treated as a communication primitive. We present SparSP, an efficient sparse sequence parallel communication system that co-designs token distribution, communication routing, and asynchronous execution for sparse video diffusion models. First, Dependency-Aware Placement distributes sequence blocks according to diffusion models' sparse attention patterns. Second, Demand-Directed KV Routing transfers KV blocks directly to requesting GPUs without intermediate relays. Third, a Decoupled Transfer Runtime separates communication from GPU computation to reduce resource contention and maximize effective bandwidth. Our evaluation shows that SparSP improves attention performance by 1.38- 1.5$\times$, achieves an average 1.17$\times$ (up to 1.69$\times$) end-to-end speedup across three representative servers and three video diffusion models, and reduces communication volume by 12.54-23.05%. Moreover, we achieve an average 1.53-1.76$\times$ bandwidth improvement over NCCL.