arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于稳定语言模型预训练的UE5M3 FP4块缩放

UE5M3 FP4 Block Scaling for Stable Language Model Pretraining

Robert Hu, Carlo Luschi, Paul Balanca

arXiv 2609.02846首次发表:更新:

AI 中文总结

本研究提出UE5M3 FP4块缩放方案,预训练Nemotron-H 8B模型,相比NVIDIA Transformer Engine v{}方案训练损失更低,吞吐量提升21.2%,推动UE{}块缩放的原生支持。

AI 中文摘要

稳定的4位浮点数(FP4)预训练难度较大,因为E2M1有效载荷仅表示狭窄的幅值范围。NVIDIA的Transformer Engine v{}方案通过电流张量缩放、随机Hadamard变换(RHT)和bfloat16(BF16)最终层解决该问题,但需在FP4矩阵乘法之外增加工作量。我们将E2M1有效载荷与无符号E5M3(UE{})块缩放配对,其更宽的幅值范围允许周期性张量缩放;我们的方案对反向梯度应用选择性随机舍入,省略RHT,并在所有符合条件的内部线性层中使用FP4。我们预训练了一个Nemotron-H 8B模型,处理近1900亿个token。与Transformer Engine v{}相比,所提出的块16方案完成时的最终窗口训练损失更低,且在各自的量化推理策略下,以保留的负对数似然衡量的验证损失更低;其量化推理下游点估计在所有三个报告的聚合指标上也更高。一项原生v{}执行 ablation(消融实验)同时移除RHT和BF16最终块豁免,使模型主体的token吞吐量提高了21.2%。这些结果证明了采用更简单方案的端到端软件模拟UEFP预训练的可行性,并推动对UE{}块缩放的原生支持。

英文摘要

Stable 4-bit floating-point (FP4) pretraining is difficult because the E2M1 payload represents only a narrow range of magnitudes. NVIDIA's Transformer Engine \nv{} recipe addresses this with current-tensor scaling, a randomized Hadamard transform (RHT), and bfloat16 (BF16) final layers, adding work outside the FP4 matrix multiplications. We instead pair E2M1 payloads with unsigned E5M3 (\ue{}) block scales. Their wider range permits periodic tensor scaling, while our recipe applies selective stochastic rounding to backward gradients, omits RHT, and uses FP4 in all eligible internal linears. We pretrain a Nemotron-H 8B model for nearly 190 billion tokens. Compared with Transformer Engine \nv{}, the proposed block-16 recipe finishes with lower final-window training loss and, under their respective quantized-inference policies, lower validation loss measured as held-out negative log-likelihood. Its quantized-inference downstream point estimates are also higher on all three reported aggregates. A native \nv{} execution ablation that jointly removes RHT and the BF16 final-block exemption increases measured model-body token throughput by 21.2\%. These results demonstrate end-to-end software-emulated \uefp{} pretraining with a simpler recipe and motivate native support for \ue{} block scaling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑