发表机构
Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出首个端到端FP4 RL后训练方案,通过HiFloat4格式与Rollout-ResQ技术,将FP4 RL与BF16的准确率差距大幅缩小,使全量化FP4 RL接近全精度水平。
AI 中文摘要
据我们所知,本文提出了首个端到端FP4强化学习(RL)后训练方案,其中包括rollout策略和训练策略(及其前向与反向传播)均以4位精度运行。系统研究表明,FP4 RL中性能下降的主要来源并非训练侧量化误差,而是rollout激活量化:异常值会将动态范围拉伸至极大,导致大量激活值在FP4格式下下溢为零。与直觉相反,将训练策略恢复为更高精度、同时保持rollout为FP4精度的做法,会使准确率比全FP4基线更差,这表明rollout与训练的不匹配是主要失效模式,排除了标准预训练式修复方案。我们提出Rollout残差量化(Rollout-ResQ)来解决该问题:这是一个仅添加到FP4 rollout矩阵乘法的、受硬件友好稀疏性模式约束的单一残差校正项,该轻量校正可恢复因异常值驱动的下溢损失的大部分精度,且不会增加rollout的计算占用。在Qwen2.5-3B和Qwen2.5-Math-7B模型上,Rollout-ResQ与HiFloat4(HiF4)格式(其三级分层缩放在FP4的严格4位预算下保留了分辨率)相结合,将与BF16的准确率差距从4.9%缩小至1.1%,使全量化FP4 RL接近全精度水平。应用于开放标准MXFP4时,相同方案将差距从13.6%缩小至5.3%,表明FP4格式选择是决定可恢复准确率上限的关键因素。这些结果共同确立了HiF4作为端到端FP4 RL后训练的支撑格式,以及Rollout-ResQ作为可缩小与BF16差距的激活侧机制。
英文摘要
We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. A systematic study reveals that the dominant source of degradation in FP4 RL is not training-side quantization error but rollout activation quantization: outliers stretch the dynamic range so far that a large number of activation values underflow to zero under FP4. Counterintuitively, restoring the training policy to higher precision while keeping the rollout in FP4 makes accuracy worse than full FP4 baseline, exposing rollout-training mismatch as the principal failure mode and ruling out standard pretraining-style fixes. We address this with Rollout Residual Quantization (Rollout-ResQ): a single residual correction term constrained to a hardware-friendly sparsity pattern, added only to the FP4 rollout matmul -- a lightweight correction that recovers most of the precision lost to outlier-driven underflow without inflating the rollout's compute footprint. On Qwen2.5-3B and Qwen2.5-Math-7B, Rollout-ResQ paired with the HiFloat4 (HiF4) format -- whose three-level hierarchical scaling preserves resolution under FP4's tight 4-bit budget -- closes the accuracy gap to BF16 from 4.9% to 1.1%, bringing fully quantized FP4 RL within striking distance of full precision. Applied to the open-standard MXFP4, the same recipe narrows the gap from 13.6% to 5.3%, revealing that FP4 format choice is a key factor that determines the ceiling on recoverable accuracy. Together, these results establish HiF4 as the enabling format for end-to-end FP4 RL post-training, and Rollout-ResQ as the activation-side mechanism that makes the gap to BF16 closable.