arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16937cs.LGcs.AIcs.PL

超越Token局部模仿:面向在线策略蒸馏的奖励兼容时间信用分配

Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation

  • School of Vehicle and Mobility & College of AI, Tsinghua University(清华大学车辆与运载学院与人工智能学院)
  • Didi Voyager Labs, DiDi Autonomous Driving(滴滴自动驾驶沃芽实验室)

机构由 AI 辅助整理,请以论文原文为准。

Shiqi Liu, Zeyu He, Letian Tao, Guojian Zhan, Jiaxin Gao, Feihong Zhang, Jingliang Duan, Wei Xiong, Kehua Sheng, Bo Zhang, Yang Guan, Shengbo Eben Li

AI总结:

针对在线策略蒸馏中目标保真度与稳定性权衡,提出γOPD方法,通过折扣时间信用分配和奖励兼容有界混合机制,在数学与代码推理任务上超越现有方法。

AI中文摘要:

在线策略蒸馏(OPD)已成为大语言模型后期训练的有效方法,然而现有目标在目标保真度与优化稳定性之间存在权衡。Token级OPD提供稳定但局部的监督,而序列级OPD以依赖于视界(horizon)的方差为代价捕获未来信用。我们建立了这些公式的统一时间信用视图,表明实用的Token级OPD可被解释为序列级反向KL梯度的时序近似。基于此联系,我们提出γOPD,其使用折扣时间信用分配来平衡长视界监督与优化稳定性,同时允许视界无关的方差界。我们进一步为γOPD开发了一种奖励兼容的有界混合(RBM)机制,该机制平衡可验证的结果反馈与折扣OPD优势,以超越纯教师依赖的优化。在数学和代码推理上的实验表明,在标准、规模不匹配和多教师蒸馏设置中,相较于现有OPD方法,该方法均有一致的改进。

英文摘要:

On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost of horizon-dependent variance. We establish a unified temporal-credit view of these formulations, showing that practical token-level OPD can be interpreted as a temporal approximation to the sequence-level reverse-KL gradient. Building on this connection, we propose $γ$OPD (GammaOPD), which uses discounted temporal credit assignment to balance long-horizon supervision and optimization stability, while admitting a horizon-independent variance bound. We further develop a reward-compatible bounded mixing (RBM) mechanism for $γ$OPD that balances verifiable outcome feedback with the discounted OPD advantage to move beyond purely teacher-dependent optimization. Experiments on mathematical and code reasoning demonstrate consistent improvements over existing OPD methods across vanilla, size-mismatched, and multi-teacher distillation settings.

↑