arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越模仿:通过推理进度过滤在线策略蒸馏

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

Chen Yang, Haiyuan Wan, Rengrong Xiong, Yize Chen, Danny H. K. Tsang

arXiv 2608.19408首次发表:更新:

AI 中文总结

针对在线策略蒸馏中教师奖励与真实推理进度不匹配的问题,提出R2-OPD方法,通过过滤冲突奖励提升推理性能,实现对标准OPD的改进。

AI 中文摘要

在线策略蒸馏(OPD)作为一种后训练语言模型的有效框架应运而生,它将学生生成的轨迹与教师提供的密集标记级监督配对。然而,OPD隐含假设教师导出的奖励是推理进度的合适替代指标,因此在策略优化过程中平等对待所有教师反馈。但在实践中,这一假设并不总是成立。我们观察到,教师导出的奖励常与真实推理进度冲突,因为具有明确推理推进的推理步骤可能仅因偏离教师输出而获得更低的蒸馏奖励。为解决这种不匹配,我们提出了面向在线策略蒸馏的推理进度感知奖励过滤(R2-OPD),它构建了轨迹内推理跨度的两个排名,一个来自教师导出的奖励,另一个来自独立估计的进度奖励。当两个排名不一致时,我们会选择性抑制蒸馏奖励,减少与推理进度冲突的监督,同时保留有效的教师指导。我们的方法相较于标准OPD表现出持续的提升,尤其在推理性能方面。

英文摘要

On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher. However, OPD implicitly assumes that teacher-derived rewards are an appropriate proxy for reasoning progress, and therefore treats all teacher feedback equally during policy optimization. While in practice, this assumption does not always hold. We observe that teacher-derived rewards often conflict with genuine reasoning progress, as reasoning steps with clear reasoning advancement may still receive lower distillation rewards, simply due to deviation from teacher's outputs. To address this mismatch, we propose Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD), which constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and the other from independently estimated progress reward. Distillation rewards are selectively suppressed whenever the two rankings disagree, reducing supervision that conflicts with reasoning progress while preserving effective teacher guidance. Our approach shows consistent improvement over standard OPD especially regarding reasoning performances.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑