AI 中文总结
针对RLVR探索不足问题,提出PSRA方法,通过贝叶斯顺序分配在无引导与策略条件提示间切换,保留策略路径,提升推理性能并减少饱和。
AI 中文摘要
基于可验证奖励的强化学习(RLVR)常因探索不足而受限:难题可能产生均匀错误的轨迹组,从而几乎没有学习信号。我们表明,此类失败未必反映能力缺失。相反,有限采样往往集中于问题特定的主导推理策略,而忽略了模型已支持的其他策略。此外,这些策略的可访问性在RL过程中会演变:有些被内化为自主行为,而另一些在被吸收前变得难以引出。受这些观察启发,我们提出问题-策略轨迹分配(PSRA),将无引导和策略条件提示视为竞争性探索臂,并使用贝叶斯顺序分配将固定轨迹预算导向最可能产生信息丰富、非饱和组的臂。一个保留目标保持有用的策略条件路径可访问,同时成功的引导行为被转移到无引导策略。在Qwen2.5模型从1.5B到7B以及两个RL训练语料库上,PSRA持续提升推理性能,减少死饱和,增强分布外迁移,并在增加推理预算下保持更大收益。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a problem-specific dominant reasoning strategy while leaving alternative strategies already supported by the model unexplored. Moreover, the accessibility of these strategies evolves during RL: some are internalized into autonomous behavior, while others become difficult to elicit before being absorbed. Motivated by these observations, we introduce Problem--Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups. A preservation objective keeps useful strategy-conditioned routes accessible while successful guided behaviors are transferred to the unguided policy. Across Qwen2.5 models from 1.5B to 7B and two RL training corpora, PSRA consistently improves reasoning performance, reduces dead saturation, strengthens out-of-distribution transfer, and maintains larger gains under increased inference budgets.
Comments23 pages, 5 figures