arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GrowMTP:强化学习能否自行长出草稿头?

GrowMTP: Can RL Grow Its Own Draft Head?

Minghua He, Lingzhe Zhang, Yuan Liu, Xiao Zhou, Aiwei Liu

arXiv 2609.16648首次发表:更新:

发表机构

WeChat AI, Tencent; Peking University(腾讯微信人工智能实验室; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对强化学习自回归生成瓶颈,提出GrowMTP,利用RL训练中rollout分布窄及验证监督信号,在RL循环内从零训练草稿头,实现最高2.13倍rollout加速。

AI 中文摘要

强化学习(RL)后训练驱动了大语言模型的前沿能力,其挂钟时间主要被自回归 rollout 生成所占据。投机解码是解决这一瓶颈的既有方法,但现有的草稿头在 RL 之前必须经过预训练或预热,从而在待加速的 RL 运行之外引入了可观的训练成本。我们观察到,RL 训练本身即提供了在线草稿头训练所需的两个条件:其 rollout 分布远窄于预训练分布,且其验证步骤持续产生与该分布对齐的监督信号。基于这些观察,我们提出 GrowMTP,它利用这些监督信号在 RL 循环内从头训练草稿头,所有头更新均与策略主干解耦。在 Qwen3-4B(无草稿头)、MiMo-7B-SFT(弱草稿头)和 Qwen3.5-4B-Base(强草稿头)上,GrowMTP 分别实现了 2.13 倍、1.93 倍和 1.36 倍的 rollout 加速,以及 1.60 倍、1.41 倍和 1.20 倍的端到端加速。因此,GrowMTP 可作为模块化组件服务于现有 RL 训练框架,尤其为没有预训练草稿头的模型提供了一条从零开始的加速路径。

英文摘要

Reinforcement learning (RL) post-training drives the frontier capabilities of large language models, with its wall-clock dominated by autoregressive rollout generation. Speculative decoding is an established remedy for this bottleneck, but existing draft heads must be pretrained or warmed up before RL, introducing substantial training cost outside the RL run to be accelerated. We observe that RL training itself provides both conditions required for online draft-head training: its rollout distribution is far narrower than that of pretraining, and its verification step continuously produces supervision signals aligned with this distribution. Building on these observations, we propose GrowMTP, which uses this supervision to train a draft head from scratch entirely within the RL loop, with all head updates detached from the policy backbone. On Qwen3-4B (no draft head), MiMo-7B-SFT (weak head), and Qwen3.5-4B-Base (strong head), GrowMTP achieves rollout speedups of 2.13x, 1.93x, and 1.36x, and end-to-end speedups of 1.60x, 1.41x, and 1.20x, respectively. GrowMTP therefore serves existing RL training frameworks as a modular component, particularly offering a from-scratch acceleration path for models without pretrained draft heads.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑