发表机构
KlingAI Research; Tsinghua University; NVIDIA(科灵人工智能研究院; 清华大学; 英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究分布式扩散Transformer推理中通信-计算重叠问题,识别出X-Stage阶段,用轻量级突发间隙模型预测相关开销等,重新设计两个通信-计算融合内核,实现加速,确立发出后进度为可测量调度杠杆。
AI 中文摘要
细粒度的设备发起通信使分布式扩散Transformer(DiT)推理中的持久GPU内核能够发出远程存储指令,并使数据移动与张量核心计算重叠。现有系统调度通信发出时间和接收数据何时可消费,但忽略了远程可见完成前的发出后进度,使发送方背压难以预测。我们识别出X-Stage,一个软件可见的发出后流水线阶段。在具有最新NVIDIA架构的八GPU节点上的测量表明,短远程存储突发会随着发出者恢复工作而耗尽,而持续注入会耗尽有限的未完成容量并延迟后续发出。一个由无背压发出时间、有效耗尽率和未完成容量参数化的轻量级突发间隙模型预测发出开销、突发间恢复和背压开始。在该模型指导下,我们重新设计了两个通信-计算融合内核。对于DeepGEMM MegaMoE,在专家波之间交错Linear-1和Linear-2工作,使计算置于集中的远程存储突发之间,在84种配置上比专家波基线实现了1.18倍几何平均和1.62倍最大内核加速。对于尤利西斯序列并行注意力,将注意力后全对全与FlashAttention进行瓦片粒度融合,使输出瓦片所有者能够发出远程存储指令并恢复计算,而无需专用通信线程或流式多处理器。FlashAttention-3和FlashAttention-4比串行执行实现了1.43倍和1.42倍的最大发送方可见加速,在长序列时其稳态时间接近单独的FlashAttention。这些结果将发出后进度确立为一个可测量的调度杠杆,用于塑造突发、避免背压和隐藏发送方开销。
英文摘要
Fine-grained, device-initiated communication allows fused GPU kernels to issue remote stores directly from their compute pipelines, a pattern increasingly used in expert parallelism (EP), tensor parallelism (TP), and Ulysses-style sequence parallelism (UP). Existing designs reason about where communication is issued and when remote data becomes ready, but lack a quantitative model of the sender-side interval after a remote store is accepted and before it becomes visible at the destination. This interval determines whether communication remains decoupled from computation or backpressures it. We identify X-Stage, a software-visible post-issue stage with finite decoupling. Downstream pressure can dissipate while the issuer resumes useful work, whereas sustained injection consumes X-Stage headroom and eventually stalls the compute pipeline. We characterize this behavior and build a calibrated model that predicts whether remote-store arrivals accumulate backpressure or recover during intervening computation. Guided by the model, we reshape bursty arrivals when they would exhaust X-Stage headroom and exploit natural compute windows when headroom can recover concurrently. Evaluation across representative EP, TP, and UP workloads shows up to 1.62x fused-kernel, 1.75x end-to-end, and 1.43x sender-visible speedup, respectively. Microbenchmarks further validate the model's predictions of backlog accumulation, recovery, and sender-side backpressure.