arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向在线蒸馏的奖励对齐重加权

Reward-Aligned Reweighting for On-Policy Distillation

Haofeng Xu, Junwei Su, Lansong Diao, Wenchao Zhou, Chuan Wu

arXiv 2609.35517首次发表:更新:

发表机构

The University of Hong Kong; Alibaba Group; University of Science and Technology of China(香港大学; 阿里巴巴集团; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出R²-OPD方法,通过结果一致性和教师-学生分歧动态重分配教师监督,解决在线蒸馏中均匀加权导致的决策效用不匹配,在数学推理和代码生成任务上显著提升学生模型性能。

AI 中文摘要

在线蒸馏(On-policy distillation, OPD)通过更强的教师模型对学生生成轨迹提供密集反馈来训练学生语言模型。然而,标准OPD对词级蒸馏项进行均匀加权,隐式地将局部教师偏好视为修正效用的代理。然而,一个决策的任务价值取决于学生如何完成后续推理。这种不匹配可能导致模仿抑制可行的学生策略,或强化学生无法可靠执行的路径。验证过的轨迹结果提供了关于延续质量的补充证据,但不能直接识别单个决策的效用。我们提出了面向在线蒸馏的奖励对齐重加权(Reward-Aligned Reweighting for On-Policy Distillation, R$^{2}$-OPD),该方法利用结果一致性和教师-学生分歧的大小来持续重新分配教师监督。它赋予奖励对齐的修正更大的相对影响,同时保留密集反馈,超越了均匀模仿和硬过滤。我们的分析形式化了局部教师偏好与学生延续价值之间的不匹配,并建立了重新分配以改善相对于均匀OPD的一阶任务进展的充分条件。在七个数学推理基准上,R$^{2}$-OPD在跨规模和同规模蒸馏中均取得了比较训练方法中最高的平均准确率。它在所有七个基准上均优于标准OPD,对于1.7B和4B学生分别平均提升3.5和2.4个百分点。扩展到代码生成任务,相较于标准OPD平均提升1.6个百分点。这些结果突显了结果引导的监督分配作为一种有效方式,能够将密集教师反馈转化为跨模型规模和任务领域的更强学生表现。

英文摘要

On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R$^{2}$-OPD), which uses outcome agreement and the magnitude of teacher--student disagreement to continuously reallocate teacher supervision. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R$^{2}$-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of 3.5 and 2.4 percentage points for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of 1.6 percentage points over standard OPD. These results highlight outcome-guided supervision allocation as an effective way to translate dense teacher feedback into stronger student performance across model scales and task domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑