arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29188cs.LGcs.AIcs.CL

锁在入口,开放在内:RLVR 如何缩小解空间

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

  • School of Future Technology, Shanghai University(上海大学未来技术学院)
  • School of Computer Science, University of Birmingham(伯明翰大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

Qiancheng Zhou, Ruizhe Li

AI总结:

该研究针对 RLVR 导致解空间收缩的问题,发现收缩集中在推理入口处,通过入口干预可提升解覆盖率且不损失准确率,早期步骤熵崩溃并非必然。

AI中文摘要:

带可验证奖励的强化学习(RLVR)大幅提升了单样本准确率(pass@1),但会导致策略的解空间收缩,降低测试时扩展的收益。本研究探究推理轨迹中这种广度损失发生在何处:是策略无法访问有效的解族,还是启动后无法执行计算?为区分访问与执行,我们分析了 Countdown 任务,其解空间可被穷尽枚举为由第一个操作数和运算符定义的离散入口族,实验采用 Qwen2.5-3B 上的 PPO 和 Qwen2.5-3B-Instruct 上的 GRPO 两种训练设置。两种设置下,解覆盖率最高下降 67%,即使是所有检查点都能解决的问题,覆盖率也减半。我们表明这种收缩主要集中在入口处:第一个算术运算前的逐词似然变化幅度是后续推理过程的 11 至 16 倍。仅提供未被选中的入口前缀,就能使低访问族的完成率提升一个数量级以上(PPO 下从 0.018 升至 0.212),证明替代方案仍可执行但不再被启动。基于此定位,我们发现表面提示无法恢复多样性,但针对入口的干预措施可行:使用早期检查点进行晚层参数插值,在不损失 pass@1 的情况下将解覆盖率提升 37%。最后,我们在 6 个数学基准上验证了 7B 和 14B 模型中早期步骤熵崩溃会重复出现,但这并非推理优化的必然副产品:SFT 基线保留了两倍以上的覆盖率,而分阶段 SFT--DPO--RLVR 流程则保留了早期步骤的熵。综上,推理广度损失在入口处,而非内部。代码:this https URL。

英文摘要:

Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes the policy's solution space to contract, diminishing the returns of test-time scaling. In this work, we investigate where inside a reasoning trajectory this breadth is lost: does the policy fail to access a valid solution family, or does it fail to execute computation once initiated? To disentangle access from execution, we analyze the Countdown task, whose solution space can be exhaustively enumerated into discrete entrance families defined by the first operand and operator, across PPO on Qwen2.5-3B and GRPO on Qwen2.5-3B-Instruct. Across both training setups, solution coverage falls by up to 67%, halving even on problems solved across all checkpoints. We show that this contraction is heavily concentrated at the entrance: per-token likelihood shifts are 11x--16x larger prior to the first arithmetic operation than during downstream reasoning. Supplying only an unselected entrance prefix restores completion rates in low-access families by over an order of magnitude (0.018 -> 0.212 under PPO), demonstrating that alternative solutions remain executable but are no longer initiated. Guided by this localization, we find that while surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1. Finally, we show that early-step entropy collapse recurs across six math benchmarks with 7B and 14B models, but is not an inevitable byproduct of reasoning optimization: an SFT baseline preserves more than double the coverage, and staged SFT--DPO--RLVR pipelines retain early-step entropy. In summary, reasoning breadth is lost at the door, not inside the room. Code: https://github.com/ershiyidian/early-branch-locking.

补充信息

↑