arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向大语言模型全流水线FP8强化学习

Towards Full Pipeline FP8 Reinforcement Learning for LLMs

Fanchao Chen, Ziheng Jiang, Ziyun Wei, Zheng Zhong, Du Li, Chi Zhang, Haibin Lin, Shivaram Venkataraman

arXiv 2609.22870首次发表:更新:

发表机构

University of Wisconsin Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对全流水线FP8强化学习训练不稳定性问题,提出校准裁剪动态方法对齐FP8裁剪边界与BF16分布,消除熵激增并恢复性能。

AI 中文摘要

强化学习(RL)已成为提升大语言模型(LLMs)推理与智能体能力的关键技术。尽管FP8量化能够加速RL训练,但在整个FP8 RL流水线中保持稳定性仍具挑战。以往工作侧重于使用如TIS等校正技术解决训练-推理不匹配问题,而我们揭示全流水线FP8 RL仍遭受严重的训练不稳定性,表现为训练中期异常熵激增和输出乱码。我们将此不稳定性追溯至一个此前被忽视的原因:复合FP8量化噪声扭曲重要性比率,不成比例地将负优势token推出信任区域,并错误地将其梯度置零。结果,病态输出未得到适当惩罚,并在训练过程中累积。为解决此问题,我们提出校准裁剪(Calibrated Clipping),一种动态方法,通过匹配下界裁剪分位数并相应重新平衡上界,使FP8裁剪边界与高精度BF16分布对齐。跨GRPO和DAPO算法、8B至32B模型规模及多种FP8缩放粒度的广泛实验表明,我们的方法成功消除熵激增,并恢复与BF16基线相当的性能。

英文摘要

Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused on resolving train-inference mismatches using correction techniques like TIS, we reveal that full-pipeline FP8 RL still suffers from severe training instability, manifesting as anomalous mid-training entropy surges and garbled outputs. We trace this instability to a previously overlooked cause: compounded FP8 quantization noise distorts the importance ratio, disproportionately pushing negative-advantage tokens outside the trust region and erroneously zeroing out their gradients. As a result, pathological outputs are not properly penalized and accumulate over the course of training. To address this, we propose Calibrated Clipping, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly. Extensive experiments across GRPO and DAPO algorithms, model scales from 8B to 32B, and multiple FP8 scaling granularities demonstrate that our approach successfully eliminates entropy surges and restores performance comparable to the BF16 baseline.

Comments17 pages, 16 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑