发表机构
University of Illinois Urbana-Champaign; Google(伊利诺伊大学厄巴纳-香槟分校; 谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出基于规则指纹选择多样化推理路线的SFT数据,显著提升后RL泛化能力,在多个基准上验证了有效性。
AI 中文摘要
验证过的解决方案对于为强化学习(RL)准备推理模型并非同等有用。我们对路线多样性(即监督微调(SFT)数据中推理步骤序列的变化)进行了全面研究,并提出了一种轻量级、基于规则的指纹来选择这种多样性。在相同预算的一个池中,采用匹配的训练配方和检查点,选择多样化而非相似的路线,能够提升后RL在谜题和数学问题上的问题覆盖率,包括那些比任一训练阶段所见问题更难的问题。在合成实验中,路线多样化的SFT使OLMo3-7B在SFT未涉及的环境中的pass@8提升了16.9个百分点。在单模型条件下,即一个模型编写每个候选,多样化选择在10个数学基准上的平均pass@8最多提升6.2个百分点。前RL诊断揭示了原因:多样化的SFT可以在更多提示上产生成功和失败的尝试,尽管平均准确率略低,这为群体相对RL提供了更多带有学习信号的提示。在3个开源语料库上,我们仅使用CPU的选择器,无需模型调用,在每次平均后RL性能比较中都优于更昂贵的替代方案。这些结果将推理路线多样性确定为选择SFT数据的实用标准,这些数据能更好地为RL准备模型。
英文摘要
Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to select for it. From one pool at one budget, with matched training recipes and checkpoints, selecting diverse rather than similar routes improves post-RL problem coverage across puzzles and mathematics, including on problems harder than those seen in either training stage. In synthetic experiments, route-diverse SFT improves OLMo3-7B's pass@8 by 16.9 points on environments held out from SFT. In a single-model condition, where one model writes every candidate, diverse selection gains up to 6.2 points of mean pass@8 across 10 mathematics benchmarks. Pre-RL diagnostics suggest why: diverse SFT can produce both successful and failed attempts on more prompts despite slightly lower mean accuracy, giving group-relative RL more prompts with a learning signal. On 3 open-source corpora, our CPU-only selector, without model calls, outperforms more expensive alternatives in every comparison of mean post-RL performance. These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL.