arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TRACE:面向MoE语言模型FP4强化学习的 rollout 引导量化感知训练

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Xin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao, Zheng Li, Junda Feng, Yuyan Luo, Yi Zhang, Yizhong Cao, Mi Zhang, Dayiheng Liu, Jianwei Zhang

arXiv 2610.07767首次发表:更新:

发表机构

Alibaba Token Hub, Alibaba Group; Ohio State University(阿里巴巴集团阿里云Token Hub; 俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TRACE通过rollout引导的量化感知训练对齐训练与rollout路径,减少FP4 RL差异,实现MoE模型高效RL训练,性能接近BF16并获5.4倍加速。

AI 中文摘要

对大型语言模型(LLM)进行后训练强化学习(RL)在rollout生成过程中会产生大量的计算和内存开销,这促使采用低精度rollout以实现高效的RL训练。然而,现有的FP4 RL方法存在一个关键局限:它们主要在训练和rollout路径上独立地优化量化精度,而不是直接减少两条量化执行路径之间的差异。在本工作中,我们提出TRACE(通过紧凑引导实现训练-rollout量化对齐),这是一个用于MoE语言模型RL训练的FP4量化框架,解决了现有FP4 RL方法的局限性。TRACE引入了rollout引导的量化感知训练,利用rollout侧的量化结果来指导训练侧的FP4舍入决策,直接减少训练-rollout差异。此外,TRACE采用了一种高效的量化信息缓存方案,选择性地保留来自更深层的尾数和指数信息,以减少由rollout引导引入的存储和通信开销。我们在四个大规模MoE语言模型上,针对推理、编码和长时程RL任务评估了TRACE。我们的结果表明,TRACE实现了联合FP4权重/激活和FP4 KV-cache rollout,其RL性能与BF16 rollout相当,同时与BF16训练策略的事后FP4量化相比,实现了高达5.4倍的rollout加速和强大的最终FP4性能。

英文摘要

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑