arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结果引导的在线策略自蒸馏

Outcome-Guided On-Policy Self-Distillation

ZheXu Wang, Mao-Lin Luo, Yankun Hong, Zi-Hao Zhou, Bo Ye, Jian Zhao, Xialiang Tong, Min-Ling Zhang, Tong Wei

arXiv 2610.05070首次发表:更新:

发表机构

Southeast University; Zhongguancun Academy; Zhongguancun Institute of Artificial Intelligence(东南大学; 中关村学院; 中关村人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线策略自蒸馏中监督噪声和训练不稳定的问题,提出结果引导的在线策略自蒸馏(OG-OPSD),动态调整散度目标和蒸馏位置,在数学推理、多模态推理及分布外任务中显著提升性能。

AI 中文摘要

在线策略自蒸馏(OPSD)相比可验证奖励强化学习(RLVR)提供了更密集的令牌级监督和更好的计算效率。然而,这种更密集的监督可能引入大量噪声和训练不稳定性。现有改进方法通常依赖高方差的逐令牌统计,并引入额外的超参数和权衡。基于RLVR中的优势公式,我们从相同视角分析OPSD目标,并纳入结果正确性信号。我们发现,朴素OPSD对不正确轨迹施加的惩罚不足且奖励过多,因为它无论结果正确性如何都应用固定的散度目标。此外,教师监督的可靠性与轨迹结果以及沿轨迹生成的累积平均教师熵相关。基于这些观察,我们提出结果引导的在线策略自蒸馏(OG-OPSD),它根据二元结果奖励和累积平均教师熵动态调整散度目标和蒸馏位置。大量实验表明,OG-OPSD在数学推理、多模态推理和分布外任务中,在Qwen3模型1.7B、4B和8B规模以及Qwen3-VL-2B上,持续优于朴素OPSD和多个强基线。

英文摘要

On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-token statistics and introduce extra hyperparameters and trade-offs. Based on the advantage formulation in RLVR, we analyze the OPSD objective from the same perspective, incorporating outcome correctness signals. We find that vanilla OPSD imposes insufficient penalties and excessive rewards on incorrect trajectories because it applies a fixed divergence objective regardless of outcome correctness. Furthermore, the reliability of teacher supervision is associated with both trajectory outcome and the cumulative average teacher entropy along the rollout. Based on these observations, we propose Outcome-Guided On-Policy Self-Distillation (OG-OPSD), which dynamically adapts both the divergence objective and distillation position according to binary outcome rewards and the cumulative average teacher entropy. Extensive experiments show that OG-OPSD consistently improves the performance of vanilla OPSD and multiple strong baselines in mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3 models at 1.7B, 4B, and 8B scales, as well as Qwen3-VL-2B.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑