发表机构
Baidu Inc.; Nanyang Technological University(百度公司; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对可验证奖励强化学习中目标模型样本语义冗余问题,提出从弱到强学习范式,引入W2SPO方法,通过注入短辅助段指导策略探索,实验表明该方法在数学推理基准测试中性能优异,能提升训练速度和准确率。
AI 中文摘要
可验证奖励的强化学习已成为增强大语言模型推理的标准方法,通常通过对比多个自生成的展开来优化策略。然而,我们发现此范式存在关键的支持有限瓶颈:在具有挑战性的推理任务上,目标模型的样本常表现出语义冗余,收敛到相同错误‘推理盆地’,为策略更新提供可忽略的奖励对比。本文提出通过从弱到强的学习范式克服此限制,用较弱但计算高效的辅助模型指导策略探索。我们引入W2SPO,一种离策略强化学习方法,将短至8个令牌的辅助段注入中间目标模型轨迹,目标模型从这些转移状态完成推理路径。基于最终可验证奖励,策略更新限于这些短插入段。实验上,W2SPO在数学推理基准测试的评估4B规模模型中表现优异,优于评估的后训练基线。与相同采样预算下的普通GRPO相比,W2SPO将Pass@1从62.3%提高到64.2%,同时实现3.55倍的训练加速。这些结果表明,弱辅助分支可通过扩展局部探索支持诱导更强的目标推理策略。
英文摘要
Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self generated rollouts. However, we identify a critical support limited bottleneck in this paradigm: on challenging reasoning tasks, the target model's samples often exhibit semantic redundancy, converging into the same erroneous "reasoning basins" that offer negligible reward contrast for policy updates. In this paper, we propose to overcome this limitation through a weak to strong learning paradigm, where a policy's exploration is informed by a weaker but computationally efficient auxiliary model. We introduce W2SPO, an off policy RL method that injects short auxiliary segments often as brief as 8 tokens into intermediate target model trajectories and the target model then completes the reasoning path from these diverted states. Policy updates are restricted to these short inserted segments based on final verifiable rewards. Empirically, W2SPO achieves superior performance among evaluated 4B scale models on mathematical reasoning benchmarks, outperforming evaluated post trained baselines. Compared with vanilla GRPO under the same sampling budget, W2SPO improves Pass@1 from 62.3% to 64.2% while achieving a 3.55 times training speedup. These results suggest that weak auxiliary branches can induce stronger target reasoning policies by expanding local exploration support.