RLVR 是核而非函数:pass@$k$ 交叉点的统计推断
RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers
浏览论文内容
中文总结 AI 辅助
本文通过统计推断揭示RLVR的pass@k交叉并非必然成立,提出用马尔可夫核建模提示条件分布,以准确预测交叉并修正功效估计,区分交叉存在与其能力含义。
中文摘要 AI 辅助
带可验证奖励的强化学习(RLVR)通常能提升 pass@1,但在更大的采样预算 $k$ 下会落后于其基础模型,这种交叉现象被视为 RLVR 仅增强了现有能力的证据。我们指出了这一解读的两个局限。首先,可见的交叉未必在统计上成立:在相同提示上比较模型时,我们构建了跨采样预算 $k$ 的置信带,这要求同时存在早期收益和后期损失的证据。在五个公开的 RLVR 配对中,初始评估均未在统计上确立交叉,而在新提示上的 32k 令牌评估定位到一次反转,首次损失出现在 11 到 61 个样本之间;功效分析表明,未能检测到交叉未必意味着不存在交叉,也解释了为何更多提示可能比每个提示更多答案更有帮助。其次,基础成功率本身并不能决定 RLVR 对提示的影响:具有相同基础成功率的提示在 RL 后具有不同的成功率,且这些差异在独立的生成半区中重复出现。这种关系是一个条件分布——即马尔可夫核——而非单一曲线,拟合该分布可预测相同提示在独立生成中的交叉,并修正简单模型的功效估计。理论进一步表明,即使训练改善了其他提示,少数最难提示上的损失也能推翻早期领先,从而将交叉存在的证据与关于其能力含义的主张区分开来。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) often improves pass@1 while falling behind its base model at larger sampling budgets $k$, a crossover read as evidence that RLVR only sharpens existing capability. We identify two limits to this reading. First, a visible crossing need not be statistically established: comparing models on the same prompts, we build confidence bands across sampling budgets $k$ that require evidence of both an early gain and a later loss. Across five public RLVR pairs no crossing is statistically established in the initial evaluations, while a 32k-token evaluation on fresh prompts locates a reversal with first loss between 11 and 61 samples; power analysis shows why failure to detect a crossing need not mean no crossing, and why more prompts can help more than more answers per prompt. Second, base success alone does not determine what RLVR does to a prompt: prompts with the same base success rate have different post-RL success rates, and these differences repeat across independent generation halves. The relationship is a conditional distribution---a Markov kernel---rather than a single curve, and fitting it predicts crossings in independent generations for the same prompts and corrects the simple model's power estimates. Theory further shows how losses on a minority of the hardest prompts can overturn an early lead even when training improves other prompts, separating evidence that a crossover exists from claims about what it means for capability.
发表机构
- Texas A&M University(德克萨斯农工大学)
- Mayo Clinic(梅奥诊所)
机构由 AI 辅助整理,请以论文原文为准。