arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dr. OPD:学习最优在线策略蒸馏中应遵循什么

Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu

arXiv 2609.38025首次发表:更新:

发表机构

Rutgers University(罗格斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线策略蒸馏中教师信号重要性不均的问题,提出加权双层优化方法Dr. OPD,通过迭代更新令牌权重与学生策略,显著提升学生性能,在数学和代码任务上超越基线。

AI 中文摘要

在线策略蒸馏(OPD)使用来自更强教师的密集、令牌级监督,在学生的自生成响应上训练学生。朴素OPD平等对待所有教师信号,假设教师的监督对每个令牌同等重要。然而,不同令牌处的教师信号对学生性能的影响可能非常不同:一些信号纠正了重要的推理错误,而另一些对最终答案影响甚微。受此观察启发,我们提出了Dr. OPD(正确完成的OPD),其定义了最优加权OPD以最大化学生性能。我们将Dr. OPD表述为一个双层优化问题,其中学生从加权教师监督中学习,而权重被选择以最大化所得学生的期望奖励。为求解Dr. OPD,我们开发了一种高效的迭代求解器,交替更新令牌权重和学生策略。在每一轮中,它以封闭形式更新权重,然后对所得加权OPD目标执行一步梯度更新。在正则性条件下,我们证明了这种加权更新比朴素OPD更新获得更高的期望奖励。在数学和代码的强到弱及同规模蒸馏实验中,Dr. OPD始终优于所有评估的基线。特别是,在强到弱蒸馏设置中,Dr. OPD将平均数学性能比朴素OPD提高了9.7个百分点,并使较小的学生能够超越其较大的教师。

英文摘要

On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by $9.7$ points over vanilla OPD, and enables the smaller student to surpass its larger teacher.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑