arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35793cs.LGstat.ML

从 Pass@K 与 Pass@1 之间的差距中学习

Learning from the Gap Between Pass@K and Pass@1

发表机构上海交通大学 · 莱斯大学 · 华东师范大学
查看机构详情
  • Shanghai Jiao Tong University(上海交通大学)
  • William Marsh Rice University(莱斯大学)
  • East China Normal University(华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

Xuan Liu, Jingbin Qian, Haosheng Chen

首次发表
浏览论文内容

中文总结 AI 辅助

提出GapFT方法,利用Pass@K与Pass@1的差距选择训练数据,在逻辑推理任务上显著提升单样本解码准确率,且数据效率更高。

中文摘要 AI 辅助

大型语言模型(LLMs)越来越多地使用可验证奖励的强化学习(RLVR)进行训练。精确的验证器还可以通过从多个样本中选择一个通过响应来支持测试时扩展,而其他部署则使用束搜索、自适应采样或工具。我们研究单样本解码,即每个查询只接收一个响应而不进行搜索,以探讨搜索暴露的行为能否被吸收到模型中。现有的验证响应后训练方法通常不区分首次解码时已解决的问题与在K个样本内恢复的失败。在固定预算下,这可能会花费示例重复部署策略已有的行为。我们引入GapFT,它根据源检查点的单样本结果选择训练证据,并针对Pass@K-Pass@1差距进行微调:即策略在单样本上失败但在K个样本内解决的问题。我们在保持目标不变的情况下匹配训练示例、处理的标记数和优化器步骤。GapFT用恢复的失败填充匹配的预算,并使用精确分解来区分恢复的失败和未恢复的失败与首次解码成功上的回归。在LogiQA 2.0和ReClor上使用Llama-3.1-8B,GapFT相比源模型将Pass@1提高了14.4和13.9个百分点,在相同学习率下优于预算匹配的均匀验证RFT,并使用三分之一的数据匹配在全验证池上微调的效果。单次解码匹配源模型验证器选择的Pass@4准确率。一项随机对照将收益归因于覆盖不同的失败,我们的分析将可用收益与可转移的失败支持联系起来。在Qwen2.5-7B上的三次重复实验保留了在两项逻辑任务上相对于均匀RFT的正收益。

英文摘要

Sampling many responses and keeping one that passes a verifier lets large language models solve problems beyond their single-response ability, but this search must be paid again for every query, while many deployments answer with a single response. Post-training on verified responses can transfer the benefit of search into the model. With a fixed budget, selecting by correctness alone spends slots on problems the model already answers correctly, leaving fewer to correct its failures. To address this imbalance, we propose GapFT, which trains on the gap between Pass@K and Pass@1: problems that the source model fails with one response but solves within K samples. GapFT keeps the objective and training budget fixed and changes only which verified responses enter training; an exact decomposition splits the resulting Pass@1 change into corrected failures and regressions on problems the source model already solved. On LogiQA 2.0 and ReClor with three model families, GapFT is above budget-matched uniform rejection-sampling fine-tuning (RFT) in every setting, with a positive pooled effect, and on Llama-3.1-8B and Mistral-7B it recovers about two thirds to four fifths of the gain of fine-tuning on the entire verified pool with 11-34% of its problems. Further analyses reveal that the gain comes from failures that the first few search samples recover, while failures found only by deeper search displace replay and add no net gain, that filling the same budget with gold-labeled failures search cannot reach lowers accuracy, and that the gain is bounded by how many transferable failures search exposes.

↑