发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过理论分析和实验证明,强化学习通过重塑策略选择而非扩展推理边界,影响大语言模型的跨领域迁移、覆盖度及计算缩放规律。
AI 中文摘要
近期关于强化学习(RL)的研究报告了关于大语言模型(LLM)推理的看似矛盾的证据。在数学上的训练可以提高其他领域的性能,然而Pass@1的提升可能与比基础模型更低的Pass@$N$同时出现。这引发了一个基本问题:RL是否扩展了LLM的推理边界,还是仅仅重新加权了其现有的推理空间?我们在Qwen和Gemma模型家族中重新审视这些现象,展示了跨领域的增益和遗忘,同时在大采样预算下的覆盖度在某些任务上增加而在其他任务上减少。对RL前后解决方案轨迹的详细分析表明,模型采用的推理策略发生了转变,这促使我们提出一个两阶段的自回归策略模型,将策略选择与问题特定的执行分开。在这个框架内,我们证明了RL的隐式偏差如何重塑策略偏好,允许在某些任务上获得增益,同时抑制其他任务所需的策略。这种机制还可以在给定采样预算下拓宽或收窄覆盖度,即使没有扩展策略支持。我们进一步为RL计算中的log-sigmoid和log-linear缩放定律提供了理论依据,并评估了它们的预测能力。总之,这些结果将策略选择的变化与跨领域迁移、推理覆盖度和计算缩放联系起来。
英文摘要
Recent studies on reinforcement learning (RL) report seemingly conflicting evidence about large language model (LLM) reasoning. Training on mathematics can improve performance in other domains, yet gains in Pass@1 can coincide with lower Pass@$N$ than the base model. This raises a fundamental question: does RL expand an LLM's reasoning boundary, or merely reweight its existing reasoning space? We revisit these phenomena across Qwen and Gemma model families, showing both cross-domain gains and forgetting, while coverage at large sampling budgets increases on some tasks and decreases on others. Detailed analysis of solution traces before and after RL indicates a shift in the reasoning strategies the model employs, motivating a two-stage autoregressive policy model that separates \emph{strategy selection} from problem-specific execution. Within this framework, we prove how RL's implicit bias reshapes strategy preferences, allowing gains on some tasks while suppressing strategies required by others. This mechanism can also broaden or narrow coverage at a given sampling budget even without expanding strategy support. We further provide theoretical justifications for log-sigmoid and log-linear scaling laws in RL compute, and evaluate their predictive power. Together, these results connect changes in strategy selection to cross-domain transfer, reasoning coverage, and compute scaling.