格式感知融合用于快速FP4预训练
Format-Aware Fusion for Fast FP4 Pretraining
浏览论文内容
中文总结 AI 辅助
提出格式感知融合方法,协同设计量化生产者与缩放域和消费者布局,在FP4预训练中实现高达37.9K tokens/s/GPU的速度,同时保持接近bfloat16的训练损失。
中文摘要 AI 辅助
四比特浮点(FP4)张量核心加速矩阵乘法,但缩放计算、操作数打包、布局构建和保存的反向状态可能抵消这一优势。我们提出格式感知融合,该方法将每个量化生产者与其缩放域及消费者布局协同设计,以支持原生MXFP、全局NVFP和协作线程数组局部NVFP。我们使用bfloat16输出投影和编译的交叉熵,评估了Llama-3系列8B模型在1600亿个token上的预训练。在匹配的同一加速器测试中,bfloat16和Transformer Engine NVFP分别达到18.8K和27.6K tokens/s/GPU,而我们最快的自定义路径达到37.9K。采用行梯度随机舍入和固定符号32值Hadamard权重梯度预条件的MXFP达到37.2K tokens/s/GPU(bfloat16模型FLOP利用率的86.3%),并以比原始bfloat16训练损失终点高2.11%结束。一个包含四个最终bfloat16块的Transformer Engine方案以27.1K tokens/s/GPU的速度,比bfloat16高0.87%。下游排名与训练损失排名不同,表明FP4结果共同依赖于缩放契约、操作数和执行路径。
英文摘要
Four-bit floating-point (FP4) Tensor Cores accelerate matrix multiplication, but scale computation, operand packing, layout construction, and saved backward state can erase the gain. We present \emph{format-aware fusion}, which co-designs each quantization producer with its scale domain and consumer layout for native \mxfp{}, global \nvfp{}, and cooperative-thread-array-local \nvfp{}. We evaluate Llama-3-family 8B pretraining through 160 billion tokens using bfloat16 output projections and compiled cross entropy. In matched same-accelerator probes, bfloat16 and Transformer Engine \nvfp{} reach 18.8K and 27.6K tokens/s/GPU, while our fastest custom route reaches 37.9K. \mxfp{} with row-gradient stochastic rounding and fixed-sign 32-value Hadamard weight-gradient preconditioning reaches 37.2K tokens/s/GPU (86.3\% bfloat16 model FLOP utilization) and ends 2.11\% above the raw bfloat16 training-loss endpoint. A Transformer Engine recipe with four final bfloat16 blocks ends 0.87\% above bfloat16 at 27.1K tokens/s/GPU. Downstream rankings differ from training-loss rankings, showing that FP4 outcomes depend jointly on scale contract, operand, and execution path.