块并行:面向高效分布式长上下文扩散语言模型训练
Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training
浏览论文内容
中文总结 AI 辅助
针对块扩散语言模型长上下文训练瓶颈,提出上下文分片块并行(CSBP),通过分片干净序列并保持损坏K/V局部化,在H200/H100上实现1.18-7.59倍吞吐量提升,并提升下游基准通过率。
中文摘要 AI 辅助
块扩散语言模型(BDLMs)将跨块的自回归依赖与块内并行去噪相结合,但长上下文训练受到分布式注意力通信和激活内存的限制。传统的上下文并行(CP)按位置对组合的干净加损坏序列进行分片,同时通信共享的干净K/V以及块特定的损坏K/V及其梯度。我们观察到BDLM目标在目标块上是可分离的。我们引入了块并行(BP),这是一种新的分布式并行维度,将每个损坏块的计算分配给一个秩(rank)。为了将BP扩展到长上下文,我们提出了上下文分片块并行(CSBP),它还将共享的干净序列在这些秩之间进行分片。CSBP保持损坏的K/V和梯度局部化,避免复制干净前缀,并保持BDLM训练语义。在16块H200 GPU上,256K上下文下,CSBP在监督微调方面比最佳基线提高了1.18-1.45倍的吞吐量,在将自回归模型转换为BDLM方面提高了1.27-1.33倍,同时匹配或降低了峰值HBM。在512K下,全模型加速达到1.61倍。在8块H100 GPU上,CSBP在512K下将DFlash2投机解码器训练加速了2.48倍,在1M下加速了7.59倍。在匹配的12小时DiffusionGemma 26B-A4B SFT运行中,CSBP在SWE-bench Verified和Terminal-Bench Lite上的每个训练检查点都实现了更高的通过率。代码:此https URL
英文摘要
Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention communication and activation memory. Conventional context parallelism (CP) shards the combined clean-plus-corrupted sequence by position, communicating shared clean K/V together with block-specific corrupted K/V and their gradients. We observe that the BDLM objective separates over target blocks. We introduce block parallelism (BP), a new distributed parallelism dimension that assigns each corrupted-block computation to one rank. To scale BP to long contexts, we introduce context-sharded block parallelism (CSBP), which also shards the shared clean sequence across those ranks. CSBP keeps corrupted K/V and gradients local, avoids replicated clean prefixes, and preserves BDLM training semantics. On 16 H200 GPUs at 256K context, CSBP improves throughput over the best baseline by 1.18-1.45x for supervised fine-tuning and 1.27-1.33x for conversion of autoregressive models to BDLMs, while matching or reducing peak HBM. Full-model speedup reaches 1.61x at 512K. On eight H100 GPUs, CSBP accelerates DFlash2 speculative-decoder training by 2.48x at 512K and 7.59x at 1M. In matched 12-hour DiffusionGemma 26B-A4B SFT runs, CSBP achieves higher pass rates at every trained checkpoint on SWE-bench Verified and Terminal-Bench Lite. Code: https://github.com/ScalingIntelligence/Turbo-dLLM
发表机构
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。