发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出探索-蒸馏(ExpDis)框架,将RLVR中的探索与优化解耦,通过训练探索者策略并蒸馏至学生策略,在不降低质量的前提下扩大探索,在数学推理基准上优于DAPO并提升多样解生成能力。
AI 中文摘要
现代语言模型在已训练好的检查点之上,通过可验证奖励的强化学习(RLVR)进行训练。RLVR的一个关键承诺是发现新的推理策略。原则上,模型可以采样其先前训练数据中不存在的新颖想法。然而,在实践中,用强新颖性激励增强RLVR的尝试成功有限,并可能降低模型质量。由于可验证奖励仅监督模型知识和行为的一小部分,这种退化难以恢复。因此,我们在一个称为“探索-蒸馏”(ExpDis)的框架中将探索与优化解耦。我们训练一个或多个探索者策略,在奖励中加入新颖性奖励,过滤其轨迹的正确性和质量,并将其蒸馏到一个单独的学生策略中。学生策略随后在不含新颖性奖励的情况下进行训练。我们重复上述过程若干轮,在探索和优化之间交替。这种解耦使我们能够在不降低学生策略质量的情况下,积极扩展探索规模。在七个数学推理基准和两个模型家族上,ExpDis在相同墙钟预算下优于DAPO。此外,我们观察到pass@$k$扩展性的改善,表明ExpDis生成的模型能产生更多样化的正确解决方案。
英文摘要
Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@$k$ scaling, indicating that ExpDis produces models that generate more diverse correct solutions.
Comments20 pages, 16 figures, 9 tables. Code: https://github.com/SaifPunjwani/Exploration-Distillation. Checkpoints: https://huggingface.co/SaifPunjwani/expdis-checkpoints