arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00220cs.LGcs.CL

可验证奖励下策略内优化中的验证器诱导支持重塑

Verifier-Induced Support Reshaping in On-Policy Optimization

Shaohang Wei, Zikun Su, Feifan Song, Wen Luo, Wei Li, Guangyue Peng, Houfeng Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现RLVR在提升当前目标时会使后续目标成功行为稀少,通过数学推理与指令遵循实验验证了验证器诱导支持重塑效应,指出端点改进无法保证策略内优化的未来可训练性

中文摘要 AI 辅助

我们表明,带有可验证奖励的策略内强化学习(RLVR)可提升当前目标,但会使后续目标的成功行为过于稀少,难以采样和强化,我们将此称为验证器诱导支持重塑,并将有效可奖励支持定义为在固定展开预算内可达的成功轨迹。我们针对两个模型族,通过重复的验证器评分采样和双向训练,在数学推理与受限指令遵循任务中研究该效应,包括使用对立验证器的顺序训练。Math-RLVR提升了平均指令遵循成功率,但在重复采样下降低了任何成功响应的提示数量。在IFEval上使用Qwen3-8B-Base时,pass@1提升6.5个百分点,而best@32下降9.8个百分点,且在两个模型和IFE基准中均出现此差异。相反,IF-RLVR使数学响应从逐步开头转向直接答案,降低了所有采样预算下的best@k,减少了后续Math-RLVR的奖励变化。令牌分布分析与受控开头干预显示,这些变化集中在前几个响应令牌中。RLVR主要对基础策略中已有的开头进行重排序,且所选开头因果影响数学可搜索性。测试的参考策略约束、路由先验和策略内蒸馏仅部分保留跨任务支持;MathIF和ReasonIF表明,边际增益仅部分转化为既正确又符合约束的响应。因此,端点改进无法保证策略内优化下的未来可训练性或联合能力。代码可在this https URL获取

英文摘要

We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce. We call this verifier-induced support reshaping and define effective rewardable support as successful trajectories reachable within a fixed rollout budget. Across two model families, we study this effect through repeated verifier-scored sampling and bidirectional training on mathematical reasoning and constrained instruction following, including sequential training with the opposite verifier. Math-RLVR raises average instruction-following success but reduces the number of prompts with any successful response under repeated sampling. On IFEval with Qwen3-8B-Base, pass@1 rises by 6.5 percentage points while best@32 falls by 9.8 percentage points, and the same divergence appears across both models and IF benchmarks. Conversely, IF-RLVR shifts math responses from step-by-step openings toward direct answers, lowers best@k across sampling budgets, and reduces reward variation for later Math-RLVR. Token-distribution analyses and controlled opening interventions show that these changes concentrate in the first few response tokens. RLVR mainly reranks openings already available in the base policy, and the selected opening causally affects math searchability. The tested reference-policy constraints, routing priors, and on-policy distillation preserve cross-task support only partially; MathIF and ReasonIF show that marginal gains translate only partly into responses that are both correct and constraint-following. Therefore, endpoint improvements do not guarantee future trainability or joint capability under on-policy optimization. Code is available at https://github.com/sylvain-wei/VISR

发表机构

  • Peking University(北京大学)
  • BUPT(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑