发表机构
Amazon; Texas A&M University(亚马逊; 德克萨斯A&M大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出APIVIS框架,将有限预算Gumbel搜索适配至分块级数学推理,通过价值引导选择与选择性监督改进RLVR训练,在多个基准上显著优于现有搜索方法。
AI 中文摘要
基于可验证奖励的强化学习(RLVR)已大幅提升大型语言模型的数学推理能力。近期研究将搜索引入RLVR的轨迹生成中,以增加轨迹多样性,但仅有多样性并不能确保搜索诱导的轨迹生成策略优于当前策略。为解决这一不足,我们提出APIVIS,一个训练时框架,将有限预算的Gumbel搜索适配到分块级数学推理。APIVIS在每个轨迹生成组内结合直接响应与搜索响应,使搜索发现的改进能产生信息性的相对奖励。它进一步对搜索改进的令牌应用选择性监督,在统一组奖励使GRPO失效时保留学习信号。我们证明,精确的价值引导选择能提高每个搜索状态的期望验证器奖励,且该保证扩展到完整轨迹生成策略,并在有界价值估计误差下具有相应的近似保证。在广泛认可的数学推理基准和不同模型规模上的实验表明,APIVIS相较于竞争性的基于搜索的方法有显著改进,验证了其有效性。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models. Recent work introduces search into RLVR rollouts to increase trajectory diversity, but diversity alone does not ensure that the search-induced rollout policy improves upon the current policy. To address this gap, we propose APIVIS, a training-time framework that adapts finite-budget Gumbel search to chunk-level mathematical reasoning. APIVIS combines direct and searched responses within each rollout group, allowing improvements found by search to produce informative relative rewards. It further applies selective supervision to search-improved tokens, preserving a learning signal when uniform group rewards render GRPO ineffective. We show that exact value-guided selection improves the expected verifier reward at each searched state and that this guarantee extends to the complete rollout policy, with a corresponding approximate guarantee under bounded value-estimation error. Experiments on widely recognized mathematical reasoning benchmarks and different model scales demonstrate substantial improvements over competitive search-based methods, validating the effectiveness of APIVIS.