SpecRoll:用于投机强化学习rollout的快慢校验器反馈自适应方法
SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts
- VNU University of Engineering and Technology(越南国家大学河内理工大学)
- Viettel AI(越南电信人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
SpecRoll是一种投机强化学习rollout引擎,通过快慢双时间尺度自适应结合相关校验机制,在5个1.5B-14B模型和3个数学推理数据集上,相比普通GRPO和FastGRPO实现了生成与端到端速度的显著提升。
中文摘要 AI 辅助
强化学习(RL)训练后可提升大语言模型的推理能力,但自回归rollout生成仍是主要效率瓶颈。投机解码可加速生成,但将其应用于RL存在困难,因为目标策略会持续演化:静态提议者会失效,而频繁的草稿模型更新会带来大量开销。我们提出SpecRoll,一种投机rollout引擎,可在双时间尺度上自适应,同时保留目标模型的采样分布。轻量未来token头生成并行提议,我们提出的Reflex模块利用延迟校验器反馈,在不进行反向传播的情况下执行有界的、轨迹局部的隐藏状态修正。互补的慢路径仅在检测到持续性能下降时才更新头参数。SpecRoll将这些机制与感知并发的稀疏树校验、精确目标校验相结合,保持目标rollout分布和GRPO目标不变。在5个参数规模从15亿到140亿的模型和3个数学推理数据集上,SpecRoll相比普通GRPO实现了1.26-2.15倍的生成加速和1.21-2.04倍的端到端加速,在所有15个匹配设置中,其生成和端到端时间均优于FastGRPO,平均成对端到端增益为1.18倍。受控 ablation 显示,快慢自适应路径提供互补收益。我们的源代码可在此https URL获取。
英文摘要
Reinforcement learning (RL) post-training improves the reasoning capabilities of large language models, but autoregressive rollout generation remains a major efficiency bottleneck. Speculative decoding can accelerate generation, yet applying it during RL is difficult because the target policy continually evolves: static proposers become stale, while frequent drafter updates add substantial overhead. We introduce SpecRoll, a speculative rollout engine that preserves the target model's sampling distribution while adapting at two timescales. Lightweight future-token heads generate parallel proposals, while our proposed Reflex module uses delayed verifier feedback to perform bounded, trajectory-local hidden-state corrections without backpropagation. A complementary slow path updates the head parameters only when sustained degradation is detected. SpecRoll combines these mechanisms with concurrency-aware sparse-tree verification and exact target verification, leaving the target rollout distribution and GRPO objective unchanged. Across five models ranging from 1.5B to 14B and three mathematical reasoning datasets, SpecRoll achieves 1.26-2.15x generation speedup and 1.21-2.04x end-to-end speedup over vanilla GRPO. It also outperforms FastGRPO in both generation and end-to-end time across all 15 matched settings, with an average pairwise end-to-end gain of 1.18x. Controlled ablations show that the fast and slow adaptation paths provide complementary benefits. Our source code is available at https://anonymous.4open.science/r/SpecRoll-26062006.