AI 中文总结
本文提出TRIAGE,一种方向感知的稳定化方法,通过片段级诊断和修复解决原生NVFP4强化学习中的失配问题,在Qwen3模型上实现全精度性能并提升2.3倍吞吐量。
AI 中文摘要
低精度执行可以大幅加速大型语言模型的强化学习(RL),但学习器与采样器执行之间的差异可能破坏策略优化的稳定性。本文刻画了失配与策略梯度方向之间的相互作用,区分了局部放大与收缩的更新贡献,这些仅凭失配幅度无法识别。在原生NVFP4运行中,我们观察到两个放大区域之间早期的不平衡,偏向于负优势、负间隙的更新。它们的尾部标记集中在响应片段的一小部分,然后失配才全局扩散。受这些发现的启发,我们提出了TRIAGE,一种方向感知的稳定化方法,利用片段级诊断选择性地重新平衡策略梯度更新,并对残余的严重失配应用有界修复。TRIAGE修改优化目标,同时在采样器和学习器上保留原生NVFP4权重和激活4位(W4A4)前向执行。在Qwen3-4B和Qwen3-30B-A3B上的实验表明,在整个评估的训练范围内优化稳定,并在五个数学推理基准上达到全精度水平性能,而带有TRIAGE的原生NVFP4相比BF16提供了高达2.3倍的rollout吞吐量。
英文摘要
Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing locally amplifying from contracting update contributions that mismatch magnitude alone cannot identify. In native NVFP4 runs, we observe an early imbalance between the two amplifying regions, favoring negative-advantage, negative-gap updates. Their tail tokens become concentrated in a small fraction of response segments before mismatch spreads globally. Motivated by these findings, we introduce TRIAGE, a direction-aware stabilization method that uses segment-level diagnosis to selectively rebalance policy-gradient updates and applies bounded repair to residual severe mismatch. TRIAGE modifies the optimization objective while retaining native NVFP4 weight-and activation 4-bit (W4A4) forward execution on both the sampler and learner. Experiments on Qwen3-4B and Qwen3-30B-A3B show stable optimization throughout the evaluated training horizon and achieve full precision level performance across five mathematical reasoning benchmarks, while native NVFP4 with TRIAGE provides up to 2.3x higher rollout throughput than BF16.