arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DualOPSD:面向在线自蒸馏的自适应特权教师

DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation

Yutong Chen, Guangfu Guo, Zhichao Xu, Kunpeng Liu

arXiv 2608.26019首次发表:更新:

发表机构

Clemson University; University of Utah(克莱姆森大学; 犹他大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出非对称交替框架DualOPSD,让在线自蒸馏的特权教师与学生策略自适应,在Qwen3-8B等模型上提升了数学推理指标并减少截断。

AI 中文摘要

在线自蒸馏(OPSD)利用学生模型的特权副本提供密集监督,无需外部教师。OPSD会固定该特权教师,即便训练期间学生分布和输出风格发生变化。我们提出DualOPSD,一种使两种策略都自适应的非对称交替框架:学生先从特权教师学习,随后教师在同一学生轨迹上向更新后的学生分布移动。该更新使后续监督能响应学习情况,且无需额外rollout。在非思考模式下的Qwen3-8B模型上,DualOPSD在AIME 2024、AIME 2025和HMMT 2025上的avg@12指标较OPSD分别提升23.61、13.89和10.00个百分点。1.7B和4B规模的结果显示,准确率提升依赖于模型规模;在所有三个规模上,DualOPSD均减少了截断,4B规模的诊断还显示教师与学生双向的KL值更低。

英文摘要

On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without an external teacher. OPSD keeps this privileged teacher fixed, even though the student distribution and output style change during training. We propose DualOPSD, an asymmetric alternating framework that adapts both policies. The student first learns from the privileged teacher. The teacher then moves toward the updated student distribution on the same student trajectory. This update makes later supervision responsive to the learner and does not require another rollout. On Qwen3-8B in non-thinking mode, DualOPSD improves avg@12 over OPSD by 23.61, 13.89, and 10.00 points on AIME 2024, AIME 2025, and HMMT 2025. Results at 1.7B and 4B show that the accuracy gain depends on model scale. Across all three scales, DualOPSD reduces truncation. The 4B diagnostic also shows lower KL in both directions between the teacher and student.

Commentspreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑