arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从基础 rollouts 到 RL 推理:预算搜索视角

From Base Rollouts to RL Reasoning: A Budgeted Search Perspective

Wenhe Sun, Cunxiang Wang, Zijun Yao, Yixin Cao

arXiv 2609.01274首次发表:更新:

发表机构

Fudan University; Zhipu AI; Tsinghua University; Shanghai Innovation Institute(复旦大学; 智谱AI; 清华大学; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究从预算搜索视角,通过统一解码框架(UDF)和预算操作点转换规则(BOPTR),探究 RL 对语言模型推理的提升本质,发现 RL 增益多源于基础模型采样效率的提升。

AI 中文摘要

带可验证奖励的强化学习(RLVR)可提升语言模型推理能力,但这些提升与推理时解码及搜索的关联仍不明确:RL 是创造了基础模型缺失的推理能力,还是将 rollout 分布转向了它本可达到却极少采样的轨迹?我们采用统一解码框架(UDF)开展行为研究,该框架将 token 级采样、类束搜索、树搜索及序列级重采样表示为共享预算操作空间上的可执行策略,事后用 pass@k、自一致性、N 选最优及首次完成成功度评分。基于 SimpleRL-Zoo 提供的配对基础/RL 检查点,我们探究 RL 默认策略曲线能否通过基础模型操作点的结构化路径近似。在 Math500、AIME、GPQA 及 IFEval 基准上,pass@k 恢复路径遵循预算操作点转换规则(BOPTR),即 $N_{\mathrm{Base}} \approx \alpha N_{\mathrm{RL}}^{\beta}$,其中指数随基准变化。在 Qwen2.5-7B 上,BOPTR 在我们测试的非最优规则中转移误差最低,为 3.41 个百分点(95% 置信区间 [2.32, 5.53]);三次随机种子重复实验结果为 3.07 ± 0.39 个百分点。该规则可扩展至四个模型家族的十个模型(拟合后添加的检查点误差为 3.28 至 4.87 个百分点)、四个未用于拟合的基准(误差 5.03 个百分点,而拟合基准误差为 4.44 个百分点),且在目标模型无 RL 检查点(误差 4.19 个百分点)或无任何 RL 监督(误差 5.08 个百分点)时仍成立。这些结果支持一种有限的内部搜索解读:在我们测试的方案下,观测到的大部分 RL 增益对应于采样效率向基础模型在搜索下本可达到的操作点的转变。我们将该缩放模式视为此方案与模型群体的描述性特征,报告其失效场景,并将 UDF 和 BOPTR 用作行为诊断工具,而非参数级等价的证据。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but how these gains relate to inference-time decoding and search remains unclear. Does RL create reasoning the base model lacks, or shift the rollout distribution toward trajectories it can already reach but rarely samples? We study this behaviorally with a Unified Decoding Framework (UDF), which expresses token-level sampling, beam-like search, tree search, and sequence-level resampling as executable policies over a shared budgeted operating space, scored post hoc with pass@$k$, self-consistency, best-of-$N$, and first-finish success. Using paired Base/RL checkpoints from SimpleRL-Zoo, we ask whether an RL default-policy curve can be approximated by a structured path of Base operating points. On Math500, AIME, GPQA, and IFEval, the pass@$k$ recovery path follows a Budgeted Operating-Point Transition Rule (BOPTR), $N_{\mathrm{Base}} \approx αN_{\mathrm{RL}}^β$, with benchmark-conditioned exponents. On Qwen2.5-7B, BOPTR gives the lowest transfer error among the non-oracle rules we test, 3.41 pp (95% CI [2.32, 5.53]); a three-seed replication gives 3.07 $\pm$ 0.39 pp. The rule extends to ten models across four families (3.28 to 4.87 pp on checkpoints added after fitting), to four benchmarks it was never fitted on (5.03 pp vs. 4.44 pp in fit), and holds without an RL checkpoint for the target model (4.19 pp) or without RL supervision of any kind (5.08 pp). These results support a qualified internalized-search reading: under the recipe we test, much of the measured RL gain corresponds to a change in sampling efficiency toward operating points the base model can already reach under search. We treat the scaling patterns as descriptive of this recipe and cohort, report where they break down, and use UDF and BOPTR as behavioral diagnostics rather than evidence of parameter-level equivalence.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑