发表机构
Shanghai Jiao Tong University; Shanghai AI Laboratory; Tsinghua University; University of Science and Technology Beijing; PJlab(上海交通大学; 上海人工智能实验室; 清华大学; 北京科技大学; PJ实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对扩散变换器并行化技术在多节点环境下通信开销大的问题,提出DiTango框架,通过选择性注意力状态机制及锚点引导的状态选择规划器等,在多节点设置中实现加速并保持生成质量。
AI 中文摘要
人工智能生成内容的最新进展推动了扩散变换器(DiTs)在高分辨率、长时间内容生成中的广泛应用。虽然并行化技术加速了扩散推理,但由于多节点环境中的过度通信开销,它们面临着重大的可扩展性挑战。我们观察到上下文并行(CP)中的序列分区表现出明显的异质性:空间上相邻的分区对注意力计算结果的贡献更大。通过将这种异构模式映射到分层通信拓扑,我们可以以降低的通信成本访问高贡献分区。这一见解激发了我们新颖的选择性注意力状态机制,该机制在去噪步骤中战略性地平衡部分注意力计算和历史结果重用。我们提出了DiTango,一种用于DiT生成的高效并行框架。DiTango具有一个锚点引导的状态选择规划器,用于优化每个分区的计算重用决策,并辅以一个运行时,用于编排高效的以状态为中心的操作。这种设计在保持生成质量的同时实现了卓越的系统效率。对流行扩散模型的实验评估表明,DiTango在多节点设置中实现了高达1.9倍的端到端和3.2倍的注意力加速,且具有近线性缩放,同时保持与最先进方法相当的生成质量。
英文摘要
Recent advances in AI-generated content have driven widespread adoption of Diffusion Transformers (DiTs) for high-resolution, long-duration content generation. While parallelization techniques accelerate diffusion inference, they face significant scalability challenges due to excessive communication overhead in multi-node environments. We observe that sequence partitions in Context Parallelism (CP) exhibit distinct heterogeneity: spatially proximate partitions contribute more significantly to attention computation results. By mapping this heterogeneous pattern to hierarchical communication topology, we can access high-contribution partitions with reduced communication cost. This insight motivates our novel selective attention state mechanism that strategically balances partial attention computation and historical result reuse across denoising steps. We present DiTango, an efficient parallel framework for DiT generation. DiTango features an anchor-guided state selection planner that optimizes computation-reuse decisions for each partition, complemented by a runtime that orchestrates efficient state-centric operations. This design achieves superior system efficiency while preserving generation quality. Experimental evaluation on popular diffusion models demonstrates that DiTango achieves up to 1.9x end-to-end and 3.2x attention speedup with near-linear scaling in multi-node settings, while maintaining generation quality comparable to state-of-the-art approaches.
Journal refThe 35th International Symposium on High-Performance Parallel and Distributed Computing (HPDC'26), July 13--16, 2026, Cleveland, OH, USA