Delta-Matching:弥合大语言模型原生8位训练的最后差距
Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
浏览论文内容
中文总结 AI 辅助
本文提出Delta-Matching方法,通过恢复softmax梯度零行和不变性,解决FP8注意力前向-后向不一致导致的优化误差,实现大语言模型原生8位训练,匹配混合精度性能。
中文摘要 AI 辅助
可靠的FP8注意力机制仍然是实现大语言模型完全原生8位训练的一大障碍。我们推导了前向-后向不一致如何产生过时的delta,并通过实验证明其如何扭曲训练动态。我们的过时delta混合运行在569M参数规模下显示出较小的损失差距,但在1.67B和5.29B规模下则出现显著的损失增加和下游性能退化。QK归一化、NoPE(无位置编码)以及较低学习率的上下文扩展可以缓解或延迟退化,但无法消除它。这一模式表明存在累积的优化误差,而较小的模型和较短的运行可以掩盖这一误差。我们提出了Delta-Matching,并证明在所述数值假设下,它能恢复softmax梯度的零行和不变性。它使得在每一个前向和后向注意力核心矩阵乘法中都能实现原生块缩放FP8,而无需改变架构、减小全局批量大小或增加辅助前向输出。在测试的各种架构、规模和训练阶段中,Delta-Matching均匹配BF16/FP32混合精度训练损失和整体下游性能。我们将发布我们的实现、训练模型和数据配方。
英文摘要
Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases and downstream degradation at 1.67B and 5.29B. QK normalization, NoPE (no positional encoding), and lower-learning-rate context extension mitigate or delay degradation without eliminating it. This pattern suggests accumulated optimization error that smaller models and short runs can conceal. We propose Delta-Matching, proving that it restores the softmax gradient's zero-row-sum invariant under the stated numerical assumptions. It enables native block-scaled FP8 in every forward and backward attention-core matmul without architectural changes, smaller global batches, or auxiliary forward outputs. Across tested architectures, scales, and training stages, Delta-Matching matches BF16/FP32 mixed-precision training loss and overall downstream performance. We will release our implementation, trained models, and data recipes.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
- NVIDIA(英伟达)
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。