arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22002cs.AI

学习随推理展开:用于高效强化学习的渐进式展开分配

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning

  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Heyang Jiang, Henry Liu, Baharan Mirzasoleiman

AI总结:

研究针对强化学习中GRPO方法计算昂贵且不稳定的问题,提出VIGOR方法,通过方差引导在线分配展开,理论上推导其加速比,实验表明该方法在数学推理和编码任务中能减少展开次数并提升通过率。

AI中文摘要:

具有可验证奖励的强化学习(RLVR)已成为改进大语言模型推理的高效框架,GRPO等方法是其最成功的实例之一。然而,GRPO依赖于重复生成长的思维链展开,训练时间与展开数量成比例,其中很大一部分是无信息的,因此计算成本高且不稳定。为缓解此问题,现有方法要么生成更大的展开池并过滤最有信息的提示,要么在训练后期利用历史信号进行过滤,这些策略虽有一定性能提升,但减慢了整体过程。为此,我们提出了方差引导在线展开分配(VIGOR),它不是为每个示例分配固定的展开预算,而是从为一批中的所有示例进行少量展开开始,迭代地将额外展开分配给组奖励方差最高的示例,直到达到固定的总展开预算。理论上,我们表明在RLVR下,奖励方差控制梯度大小,并推导了VIGOR相对于GRPO的闭式加速比,该加速比在帕累托分布的奖励方差下随细化轮数增长。在数学推理和编码任务上的实验表明,VIGOR在数学上以少达2.3倍的展开次数达到目标准确率,在编码上以少1.49倍的展开次数达到GRPO的最终全通过率,并将编码平均测试通过率提高了3.4个百分点。

英文摘要:

Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful instantiations. However, GRPO relies on repeated generation of long chain-of-thought rollouts. Training time scales with the number of rollouts, a large fraction of which are uninformative. Thus, GRPO is computationally expensive and unstable. To mitigate this, existing approaches either generate a larger pool of rollouts and filter the most informative prompts, or leverage historical signals for filtering at later stages of training. These strategies offer modest performance gains, but slow down the overall process. To address this, we propose VarIance Guided Online Rollout allocation (VIGOR) which instead of allocating a fixed rollout budget per example, begins with a small number of rollouts for all examples in a batch and iteratively allocates additional rollouts to those with the highest group reward variance until a fixed total rollout budget is reached. Theoretically, we show that under RLVR, reward variance controls the gradient magnitude, and derive VIGOR's closed-form speedup ratio over GRPO, which grows with refinement rounds under Pareto-distributed reward variance. Experiments on mathematical reasoning and coding tasks show that VIGOR reaches target accuracy with up to 2.3$\times$ fewer rollouts on math, reaches GRPO's final coding full pass rate with 1.49$\times$ fewer rollouts, and improves the coding average test pass rate by 3.4 points.

↑