arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

捕获图下低精度流水线并行训练的显式状态与资源契约

Explicit State and Resource Contracts for Low-Precision Pipeline Parallel Training under Captured Graphs

Genlang Chen, Junyi Zhu

arXiv 2609.23536首次发表:更新:

发表机构

NingboTech University; Dalian Ocean University(宁波工程学院; 大连海洋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对FP8低精度流水线训练中CUDA Graphs隐藏状态与资源管理问题,提出QEffect显式契约运行时,通过四个不变量保证正确性,实现1.82-2.79倍加速。

AI 中文摘要

CUDA Graphs通过重放静态虚拟地址上的张量操作来消除启动开销。然而,FP8流水线训练会持续改变这些固定地址所表示的缩放状态、微批次和延迟反向传播任务。拆分反向传播调度(如1F1B、Zero-Bubble)将输入梯度($dI$)和权重梯度($dW$)计算解耦以最小化气泡,打破了传统的后进先出(LIFO)生命周期。标准数据流图无法告知运行时隐藏的数值更新、非LIFO工作所有权或缓存有效性,导致静默的跨流数据损坏。我们提出QEffect,一个用于低精度流水线训练的显式状态与资源契约运行时。QEffect形式化了四个基础不变量:(1)隐藏缩放更新的时间序列化,(2)保留反向传播资源的代际所有权,(3)跨优化器边界的缓存权重有效性版本管理,以及(4)双向调用者图流完成同步。这些不变量统一管理即时执行和捕获执行,允许临时图资源在进程重启后干净地重建。利用工作所有权语义,我们还引入了一种仿射直接梯度放置机制,消除了冗余的内存复制。与TorchTitan和NVIDIA Transformer Engine集成,QEffect在延迟缩放翻转中与原生基线保持严格的位级一致性,确定性捕获跨流顺序违规,并实现无瑕的冷启动恢复。在NVIDIA H800 GPU上,捕获的Transformer层相比即时执行实现了1.82--2.79倍的加速,而直接梯度放置通过消除每个秩步96次矩阵复制额外提供了1.132倍的增益。

英文摘要

CUDA Graphs eliminate launch overheads by replaying tensor operations over static virtual addresses. However, FP8 pipeline training continuously alters the scaling states, microbatches, and deferred backward tasks that those fixed addresses represent. Split-backward schedules (e.g., 1F1B, Zero-Bubble) decouple input-gradient ($dI$) and weight-gradient ($dW$) computations to minimize bubbles, breaking traditional LIFO lifecycles. Standard dataflow graphs cannot inform the runtime of hidden numerical updates, non-LIFO work ownership, or cache validity, leading to silent cross-stream data corruption. We present QEffect, an explicit state and resource contract runtime for low-precision pipeline training. QEffect formalizes four foundational invariants: (1) temporal serialization of hidden scaling updates, (2) generational ownership of retained backward resources, (3) validity versioning for cached weights across optimizer boundaries, and (4) bidirectional caller-graph stream completion synchronization. These invariants uniformly govern eager and captured execution, allowing ephemeral graph resources to be cleanly rebuilt across process restarts. Leveraging work-ownership semantics, we also introduce an affine direct-gradient placement mechanism that eliminates redundant memory copies. Integrated with TorchTitan and NVIDIA Transformer Engine, QEffect maintains strict bitwise parity with native baselines across delayed-scaling rollovers, deterministically traps cross-stream ordering violations, and enables flawless cold-start resumption. On NVIDIA H800 GPUs, captured Transformer layers achieve 1.82--2.79x speedup over eager execution, while direct gradient placement delivers an additional 1.132x gain by eliminating 96 matrix copies per rank-step.

Comments12 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑