HyDra:在生产规模下揭秘并驯服动态上下文并行
HyDra: Demystifying and Taming Dynamic Context Parallelism at Production Scale
浏览论文内容
中文总结 AI 辅助
HyDra通过负载驱动调度和嵌套CP引擎,平衡动态上下文并行中的计算,减少流水线气泡,提升生产规模下的训练吞吐量。
中文摘要 AI 辅助
长上下文训练运行在长度跨越数量级的序列上,动态上下文并行(DCP)为每个序列分配其自身的CP度数。现有的DCP系统要么无法扩展,要么在主流水线模型上表现不佳,使得Megatron-Core(Mcore)DCP成为生产规模下的唯一选择。然而,Mcore DCP将每个度数的大小设置为适应内存,这随长度线性增长,而注意力则二次增长,因此可比的token数量掩盖了不等的计算量。在我们使用Mcore DCP进行的超过11K GPU的生产级256K上下文训练作业中,每个rank的微批量时间在相似token数量下差异高达5倍。这种偏差导致46%的流水线气泡和13%的数据并行气泡。我们提出了HyDra,一个可扩展的负载驱动DCP系统。其调度器通过将每个rank拉向一个负载目标来平衡计算,而平衡反过来使该目标可以闭式求解。它使用惰性堆而非整个池扫描来放置序列。这种平衡要求CP度数大于内存所需,因此其CP引擎在浅层外部环中嵌套内部Ulysses组,让更高的度数同时降低计算和通信。在两种规模下的评估显示了一致的增益。在512-GPU测试台上,HyDra在32K上下文下平均吞吐量比Mcore DCP提高1.18倍,在256K下提高2.48倍。在2,048-GPU生产作业中,它将流水线气泡从36%缩小到14%,将调度时间减少2.3倍,并将吞吐量比Mcore DCP提高1.10-1.43倍(平均1.25倍),比静态CP提高1.33-1.90倍(平均1.59倍)。
英文摘要
Long-context training runs on sequences whose lengths span orders of magnitude, and dynamic context parallelism (DCP) gives each sequence its own CP degree. Existing DCP systems either do not scale or perform poorly on mainstream models, leaving Megatron-Core (Mcore) DCP as the only option at production scale. Mcore DCP, however, sizes each degree to fit memory, which grows linearly with length while attention grows quadratically, so comparable token counts hide unequal computation. In our production 256K-context training job on more than 11K GPUs under Mcore DCP, per-rank microbatch times differ by up to 5x at similar token counts. The skew leads to a 46% pipeline bubble and a 13% data-parallel bubble. We present HyDra, a scalable load-driven DCP system. Its scheduler balances computation by pulling every rank toward one load target, and balance in turn makes that target solvable in closed form. It places sequences with lazy heaps rather than whole-pool scans. That balance asks for CP degrees larger than memory requires, so its CP engine nests an inner Ulysses group in a shallow outer ring, letting a higher degree lower computation and communication together. Evaluation at both scales shows consistent gains. On a 512-GPU testbed, HyDra raises throughput over Mcore DCP by 1.18x on average at 32K context and 2.48x at 256K. On a 2,048-GPU production job, it shrinks the pipeline bubble from 36% to 14%, cuts scheduling time by 2.3x, and raises throughput by 1.10-1.43x (avg. 1.25x) over Mcore DCP and 1.33-1.90x (avg. 1.59x) over static CP.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Tencent(腾讯)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。