发表机构
University of Illinois at Urbana-Champaign; Tsinghua University(伊利诺伊大学厄巴纳-香槟分校; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出策略空间扩展方法SURGE,利用强化学习历史中的检查点组合生成更强策略,无需额外训练或推理计算,在数学和编码基准上超越原生最优检查点。
AI 中文摘要
扩展推理能力通常会在强化学习(RL)或推理上投入更多计算。我们证明,一段完整的RL训练历史可以产生比其优化器所访问的检查点更强的策略。我们称之为策略空间扩展:在不延长训练或增加每次查询推理计算量的情况下,扩大从固定RL历史中可部署的策略集。我们通过SURGE(基于特征空间融合的无梯度RL扩展)实例化这一方法。SURGE结合了同一RL运行中的两个检查点:一个高精度的锚点和一个生成较短响应的竞争性供体。它将两个检查点表示为相对于其共享初始化的变化,然后对锚点的更新进行谱分解,以保留其主导分量并整合供体的互补分量。在固定锚点更新保留目标的情况下,SURGE根据权重确定块大小,而无需测试候选策略。我们评估了两个1.5B数学推理历史(DeepSeek和Nemotron)以及一个7B编码历史(OLMo)。SURGE在两个输入检查点上均提高了基准平均准确率,同时使用的推理令牌少于锚点。它在DeepSeek AIME24上达到54.17%,而实测原生最大值为50.83%;在OLMo HumanEval+上达到83.7%,而基线为82.8%。这些提升超过了观察到的训练曲线。几何控制支持RL更新结构的重要性,而不仅仅是权重位移或令牌减少。每个构建的模型作为单一策略运行。我们的发现将存储的RL历史识别为可复用的扩展资源:训练运行中可用的能力不必止步于其最佳检查点。
英文摘要
Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its dominant component and incorporate the donor's complementary component. With a fixed target for how much of the anchor update to retain, SURGE determines the block size from the weights without testing candidate policies. We evaluate two 1.5B mathematical-reasoning histories, DeepSeek and Nemotron, and one 7B coding history, OLMo. SURGE improves benchmark-average accuracy over both input checkpoints while using fewer reasoning tokens than the anchor. It reaches 54.17% on DeepSeek AIME24 against a measured native maximum of 50.83%, and 83.7% on OLMo HumanEval+ against 82.8%. These gains exceed the observed training curves. Geometric controls support the importance of RL-update structure beyond weight displacement or token reduction alone. Each constructed model runs as a single policy. Our findings identify stored RL history as a reusable scaling resource: the capability available from a training run need not end at its best checkpoint.
CommentsCorrected a typo in an author's name. No changes to the paper content