发表机构
Accenture Strategy & Consulting; Vizuara AI Labs(埃森哲战略与咨询; Vizuara AI实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示 pass@k 仅关注正确率而忽略输出多样性,通过对比 GRPO 与 RFT 训练实验,证明该指标无法检测多样性变化,且其能力提升结论因缺少基线对照而可能误导。
AI 中文摘要
pass@$k$(模型在 $k$ 次采样尝试中解决问题的比例)是该领域判断基于可验证奖励的强化学习(RL)训练后是否改进模型的默认协议。在总体层面,pass@$k$ 仅取决于问题正确样本的概率,而不考虑该概率在输出中的分布方式。我们证明这一差距并非学术问题。使用组相对策略优化(GRPO)和拒绝采样微调(RFT,即训练模型自身最短的通过验证器的轨迹)在小学数学上训练 Qwen2.5-1.5B-Instruct,使得三个互补的多样性度量(词元级熵、答案级熵、每个提示的独特答案数)朝相反方向变化,且每个分支的三个种子之间零重叠。这一差距在仅考虑验证器正确的完成结果时依然存在(在控制长度后,GRPO 的正确解词汇多样性低 15%),并且通过仅隔离错误答案多样性的计数控制检查,排除了 GRPO 更高准确率单独解释该差距的可能性。然而,pass@8 和 pass@32 在 GSM8K 上未显示一致的胜者,而一个困难的 MATH-500 子集显示相同模式:仅在低 $k$ 时出现分离。与起始检查点相比,没有训练分支显著改善难题覆盖率:RFT 显著更差,而 GRPO 在统计上与起始检查点无显著差异——因此 GRPO 在 pass@1 上对 RFT 的优势反映的是相对于基线的较小损失,而非能力增益,这是一个缺失对照的问题,而非 pass@$k$ 的失败。在 GSM8K 上,只有 pass@1(其构造上不涉及检测多样性)能清晰分离各分支,奖励了正确解多样性最低的分支。我们认为这是标准评估协议缺失其常规用于认证的属性的一具体实例。
英文摘要
pass@$k$, the fraction of problems a model solves within $k$ sampled attempts, is the field's default protocol for deciding whether reinforcement-learning (RL) post-training on verifiable rewards improved a model. At the population level, pass@$k$ depends only on a problem's probability of a correct sample, with no term for how it is distributed across outputs. We show this gap is not academic. Training Qwen2.5-1.5B-Instruct on grade-school math with Group Relative Policy Optimization (GRPO) and with rejection-sampling fine-tuning (RFT, training on the model's own shortest verifier-passed rollout) moves three complementary diversity measures (token-level entropy, answer-level entropy, unique answers per prompt) in opposite directions, with zero overlap across three seeds per arm. The gap survives restricting to verifier-correct completions only (lexical diversity among correct solutions is 15% lower for GRPO, after controlling for length) and a count-controlled check isolating diversity among incorrect answers alone, ruling out that GRPO's higher accuracy alone explains it. Yet pass@8 and pass@32 show no consistent winner on GSM8K, and a hard MATH-500 subset shows the same pattern: separation only at low $k$. Compared against the starting checkpoint, no trained arm significantly improves hard-problem coverage: RFT is significantly worse, while GRPO is statistically indistinguishable from it - so GRPO's pass@1 edge over RFT reflects a smaller loss relative to Base, not a capability gain, a missing-control issue, not a failure of pass@$k$. On GSM8K, only pass@1, with no role in detecting diversity by construction, separates the arms cleanly, rewarding the arm whose correct solutions are least diverse. We argue this is a concrete instance of a standard evaluation protocol missing a property it is routinely used to certify.
Comments10 pages, 2 figures. Accepted to the NeurIPS 2026 Workshop on Transitioning from Pre-training to Post-training (non-archival)