AI 中文总结
研究旨在解决大语言模型训练从FP8降至FP4时的不稳定性问题,提出基于2D块FP4量化的低精度训练框架,结合多种策略控制量化误差,在多个模型上评估,实现稳定的端到端FP4训练,性能与BF16相近。
AI 中文摘要
降低训练精度是提高大语言模型训练效率的关键手段,但从FP8进一步降至4位浮点(FP4)因优化过程中的不稳定性仍具挑战。我们发现现有微缩放方法中这种不稳定性的根本来源:张量转置引起的尺度不一致。在传统的1D块量化中,转置后前向和反向传播为相同值分配不同缩放因子,导致梯度更新有偏差且不稳定。为解决此问题,我们提出基于2D块FP4量化的低精度训练框架,强制转置不变缩放并保持前后向计算的一致性。我们还将其与无截断缩放和随机舍入相结合以控制量化误差并保持无偏梯度。为处理注意力机制的敏感性,我们对查询和键投影采用MXFP8量化,产生实用的混合精度设计。我们在高达7B参数的密集大语言模型和30B专家混合模型上进行评估,在多达100B令牌上训练。在所有设置中,我们的方法实现了稳定的端到端FP4训练,紧密匹配BF16性能,困惑度和下游准确率下降不到1.3%。这些结果表明,强制前后向缩放一致性足以实现大规模实用的FP4训练,为更高效的大语言模型训练提供了简单有效的途径。
英文摘要
Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging due to instability during optimization. We identify a fundamental source of this instability in existing microscaling approaches: scale inconsistency induced by tensor transposition. In conventional 1D block quantization, forward and backward passes assign di erent scaling factors to the same values after transposition, leading to biased and unstable gradient updates. To address this issue, we propose a low-precision training framework based on 2D block FP4 quantization, which enforces transposition-invariant scaling and preserves consistency between forward and backward computations. We further combine this with truncation-free scaling and stochastic rounding to control quantization error and maintain unbiased gradients. To handle the sensitivity of attention mechanisms, we adopt MXFP8 quantization for query and key projections, yielding a practical mixed-precision design. We evaluate our method on dense LLMs up to 7B parameters and a 30B Mixture-of-Experts model, trained on up to 100B tokens. Across all settings, our approach achieves stable end-to-end FP4 training and closely matches BF16 performance, with less than 1.3% degradation in perplexity and downstream accuracy. These results demonstrate that enforcing forwardbackward scaling consistency is su cient to enable practical FP4 training at scale, providing a simple and e ective pathway toward more e cient LLM training.