发表机构
Alibaba Inc(阿里巴巴公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对MoE大语言模型RL中展开生成瓶颈,提出QUADS方法。通过训练 - 推理误差分析确定激活误差是FP4 RL不稳定主因,在训练器和展开侧分别采取措施,实现BF16精度,提升指标并提高展开吞吐量。
AI 中文摘要
在用于专家混合(MoE)大语言模型的强化学习(RL)中,展开生成是一个主要瓶颈,促使诸如FP8等低精度展开加速。作为一种新兴的低精度格式,NVFP4将用于精度保持的细粒度缩放与原生W4A4 FP4通用矩阵乘法相结合,以实现比FP8更高的吞吐量。然而,直接将NVFP4应用于MoE RL展开是不切实际的。NVFP4展开与BF16训练在大约150步后崩溃,伴随着展开 - 训练器对数概率差距的迅速扩大。通过训练 - 推理误差分析和控制消融,我们确定激活误差而非权重误差是FP4 RL不稳定的主要来源。为了稳定用于MoE的NVFP4 RL,我们提出了双边量化误差对齐(QUADS)。在训练器方面,我们引入非对称量化感知训练,对权重进行伪量化,同时保持激活不量化以实现更好的对齐。在展开方面,残差激活补偿在保留原生W4A4通用矩阵乘法的同时纠正高误差激活通道。在多个基准上的MoE RL实验中,QUADS实现了BF16级别的精度,比朴素的NVFP4 RL平均pass@值提高了21.49分,并且比FP8的展开吞吐量高约16%。
英文摘要
Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8. As an emerging low-precision format, NVFP4 combines fine-grained scaling for accuracy preservation with native W4A4 FP4 GEMMs for higher throughput than FP8. However, we find that directly applying NVFP4 to MoE RL rollout is impractical. NVFP4 rollout with BF16 training collapses after roughly 150 steps, accompanied by rapidly growing rollout-trainer log-probability gaps. Through training-inference error analysis and controlled ablations, we identify activation error, rather than weight error, as the dominant source of FP4 RL instability: weights can be synchronized and aligned by a shared quantization-dequantization path, whereas activations are recomputed online and error is amplified by the coarse E2M1 grid. Therefore, to stabilize NVFP4 RL for MoE, we propose QUantization-error Alignment across Dual Sides (QUADS). On the trainer side, we introduce Asymmetric Quantization-Aware Training fake-quantizing weights while keeping activations unquantized for better alignment. On the rollout side, Residual Activation Compensation corrects high-error activation channels while preserving native W4A4 GEMMs. In our MoE RL experiments on several benchmarks, QUADS achieves BF16-level accuracy, improves average pass@1 by 21.49 points over naive NVFP4 RL, and delivers ~16% higher rollout throughput than FP8.