AI 中文总结
DuoOPD通过师生联合结果指导在线策略蒸馏,根据学生与教师成功情况动态调整反馈方向与权重,在Qwen3和Llama上平均宏精度较OPD提升2.58和5.98个百分点,并优于五个基线。
AI 中文摘要
在线策略蒸馏(OPD)使用来自更强教师的词元级反馈,在学生的自身响应上训练学生,然而教师可能会在学生已经正确回答的问题上失败,并且每个模型在不同任务上的成功频率各不相同。OPD忽略了这些结果,并且平均而言,甚至会压低学生原本正确的响应;根据学生的正确性对反馈进行门控可以修正方向,但无论教师是否成功,都以相同方式使用教师。我们提出了DuoOPD,其中学生的结果决定反馈的方向,师生联合结果决定教师如何支持该反馈:当只有教师成功时,其经过验证的答案成为评分学生失败响应的上下文;当只有学生成功时,任务内共享的权重会强化整个响应。单一规则覆盖所有四种结果组合,无需针对任务进行特定设置。在Qwen3和Llama上,DuoOPD在平均宏精度上优于所有五个基线,相比OPD分别提升了2.58和5.98个百分点,并且在另外两个涵盖科学计算、指令遵循和代码生成的任务混合上也领先。消融实验表明,仅基于结果的方向性接近门控基线,而联合结果设计贡献了大部分增益。
英文摘要
On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.