RollVerify:弥合长尾 rollout 强化学习中的效率与准确性
RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对长尾 rollout 强化学习中的 GPU 气泡和离策略样本问题,提出 RollVerify 框架,通过 OPS 度量主动验证并修复部分生成轨迹,在保持效率的同时达到与在线策略训练相当的准确性。
中文摘要 AI 辅助
强化学习对于提升大型语言模型的推理和泛化能力至关重要。它依赖于大量的 rollout,而随着上下文窗口的增长,这些 rollout 的长度变得越来越呈长尾分布。在在线策略训练中,这些长尾 rollout 可能导致 GPU 气泡,降低系统利用率并限制强化学习的可扩展性。异步或部分 rollout 方法通过放宽同步来提高吞吐量,但不可避免地会引入过时的离策略样本(轨迹),这可能损害最终准确性。现有方法主要通过重新加权训练中的离策略样本来缓解这一离策略问题,但与完全在线策略训练相比,它们仍可能留下性能差距。在这项工作中,我们不是被动地在训练中重新加权样本,而是提出 RollVerify,一个基于部分 rollout 的轻量级强化学习框架,它主动在样本进入训练之前进行验证和修复。具体来说,它引入了一个离策略偏移度量 OPS,用于量化部分生成轨迹的离策略偏差。在 OPS 约束的指导下,RollVerify 执行序列级和令牌级验证,以识别并截断轨迹的无效后缀。这产生了高质量样本,保护了模型的准确性,同时保留了部分 rollout 的效率优势。在数学和工具辅助数学推理上的实验表明,RollVerify 在降低训练成本的同时,达到了与在线策略训练相当的准确性。额外的代码生成结果提供了超越数学领域的初步证据。
英文摘要
Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models' accuracy while preserving the efficiency gains of partial rollout. Experiments on mathematical and tool-assisted mathematical reasoning show that RollVerify achieves accuracy comparable to on-policy training while reducing training cost. Additional code-generation results provide preliminary evidence beyond mathematics.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Central South University(中南大学)
- SenseTime Research(商汤科技研究院)
- The Chinese University of Hong Kong(香港中文大学)
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。