RISE:基于自外推策略蒸馏的递归改进
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
浏览论文内容
中文总结 AI 辅助
该研究提出RISE方法,通过从自身RLVR轨迹构建合成教师模型,以递归改进机制结合RLVR与OPD,在多类任务上优于仅RLVR训练及在线策略自蒸馏。
中文摘要 AI 辅助
在线策略蒸馏(OPD)为语言模型的后训练提供了密集的逐词监督,但其有效性受限于教师模型的质量:外部教师存在分布不匹配问题,而带有特权条件的自蒸馏则受限于上下文学习能力。我们提出RISE(Recursive Improvement via Self-Extrapolating Policy Distillation,基于自外推策略蒸馏的递归改进),该方法直接从模型自身的RLVR训练轨迹中构建合成教师模型。通过外推当前检查点与后续锚点之间在参数空间或输出对数几率空间的位移,RISE将稀疏的结果诱导参数更新转换为密集的词级目标,无需任何外部模型或特权条件。RISE以互补循环结合RLVR与OPD:结果奖励将外推导向正确推理,而外推教师则优化词级决策。此外,由于教师会随学生模型的改进在每次迭代中更新,蒸馏成为一种递归改进机制,而非一次性压缩步骤。涵盖数学推理、多领域STEM、代码生成及多回合智能体任务的实验表明,RISE在所有设置下均优于仅使用RLVR的训练及在线策略自蒸馏。
英文摘要
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectiveness is bottlenecked by teacher quality: external teachers suffer from distribution mismatch, while self-distillation with privileged conditioning is limited by in-context learning capacity. We propose \textbf{RISE} (\textbf{R}ecursive \textbf{I}mprovement via \textbf{S}elf-\textbf{E}xtrapolating Policy Distillation), which constructs a synthetic teacher directly from the model's own RLVR training trajectory. By extrapolating the displacement between the current checkpoint and a trailing anchor---in parameter space or output logit space---RISE converts a sparse outcome-induced parameter update into a dense token-level target, without any external model or privileged conditioning. RISE combines RLVR and OPD in a complementary loop: outcome rewards ground the extrapolation toward correct reasoning, while the extrapolated teacher refines token-level decisions. Moreover, since the teacher is refreshed every iteration as the student improves, distillation becomes a recursive improvement mechanism rather than a one-shot compression step. Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.