当可验证奖励的强化学习缩小推理边界时:诊断Pass@k反转
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
浏览论文内容
中文总结 AI 辅助
研究RLVR中的Pass@k反转问题,提出双模式解释和PBA方法,通过实验表明PBA能提升\PassK{1}和覆盖率,诊断出普通GRPO丢失基础可解决提示,而PBA保留罕见阳性轨迹,强调训练后推理要兼顾优化强度与提示安全性。
中文摘要 AI 辅助
具有可验证奖励的强化学习(RLVR)可以提高单样本准确率,但在重复采样时会使模型表现变差。我们研究了这种Pass@k反转:训练后,在大k值时策略可能比其基础模型解决的不同问题更少。失败集中在边界提示上,基础模型包含的罕见正确轨迹可通过采样恢复,但过于稀疏以至于在有限的RLVR展开组中难以可靠出现。我们认为双模式解释将此视为证据缺失失败:罕见正确轨迹可能在RLVR采样并充分强化它们之前消失。主要贡献是这种诊断和机制框架。每个问题的基础锚定(PBA)是一个故意简单的概念验证:用足够的冻结基础正确证据锐化提示,并将有风险的提示锚定到基础分布。在Omni-MATH-Test上的三个训练种子中,以MATH500作为二级高覆盖率验证基准,PBA在\PassK{1}和高预算覆盖率方面都比匹配的GRPO有所提高。一项3000提示的机制控制诊断研究在种子间一致,具有预期特征:普通GRPO会丢失基础可解决的边界提示,而PBA保留罕见的验证器阳性轨迹。我们将数学验证器用作验证器引导优化的受控测试平台;当外部工具或验证器检查重复的视觉、空间或图表推理尝试时,相同的Pass@k反转风险适用于与ECCV相关的视觉语言智能体。训练后的推理不仅应决定优化的强度,还应决定哪些提示优化是安全的。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study this pass@k inversion: after training, the policy may solve fewer distinct problems than its base model at large $k$. The failure concentrates on boundary prompts, where the base model contains rare correct trajectories that are recoverable by sampling but too sparse to reliably appear in finite RLVR rollout groups. We argue that a two-mode account explains this as an absence-of-evidence failure: rare correct trajectories may disappear before RLVR samples and reinforces them often enough. The main contribution is this diagnostic and mechanistic framing. Per-Problem Base Anchoring (PBA) is a deliberately simple proof-of-concept: sharpen prompts with sufficient frozen-base correct evidence, and anchor risky prompts to the base distribution. Across three training seeds on Omni-MATH-Test, with MATH500 as a secondary high-coverage validation benchmark, PBA improves both \PassK{1} and high-budget coverage over matched GRPO. A 3000-prompt regime-controlled diagnostic study is consistent across seeds with the expected signature: ordinary GRPO loses base-solvable boundary prompts, while PBA preserves rare verifier-positive trajectories. We use mathematical verifiers as a controlled testbed for verifier-guided optimization; the same pass@k inversion risk applies to ECCV-relevant vision-language agents when repeated visual, spatial, or chart-reasoning attempts are checked by external tools or verifiers. Reasoning post-training should decide not only how strongly to optimize, but which prompts are safe to optimize.
发表机构
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。