arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21619cs.AI

校准教师-学生差异用于同策略蒸馏

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

  • State Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
  • College of Software Engineering, Southeast University(东南大学软件工程学院)
  • Nanjing University of Posts and Telecommunications(南京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

Qiangqiang He, Jin Li, MingCai Chen

AI总结:

针对同策略蒸馏中教师自身偏差被误学的问题,提出Cal-OPD,通过正负特权干预校准差异,仅保留约52-65%信号,在数学推理基准上持续超越标准OPD。

AI中文摘要:

同策略蒸馏(OPD)通过学习更强教师与同策略学生之间的令牌级差异来改进推理模型。然而,这种差异并不纯粹反映教师与学生之间的能力差距:它还包含由教师自身引起的偏差,这些偏差因此被混入观察到的教师-学生差异中,并在标准OPD训练过程中被不加区分地学习。这一问题在特权OPD中进一步加剧,因为特权信息导致教师侧似然偏移更大,从而鼓励学生学习更多教师自身的偏差。我们引入了校准同策略蒸馏(Cal-OPD),该方法通过正向和负向特权干预来估计教师的自我偏差区域,并通过仅保留超出该区域的成分来校准原始的教师-学生差异。在数学推理基准上的实验表明,尽管仅保留原始教师-学生差异的约52%至65%作为优化信号,Cal-OPD在多种模型规模下始终优于标准OPD及其变体。

英文摘要:

On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during training. This issue is further exacerbated by privileged OPD, where privileged information induces larger teacher-side likelihood shifts, thereby encouraging the student to learn more of the teacher's own deviation. We introduce \textbf{Calibrated On-Policy Distillation (Cal-OPD)}, which estimates the teacher's self-deviation region through positive and negative privileged interventions and calibrates the original teacher--student discrepancy by retaining only the component that lies beyond this region. Experiments on mathematical reasoning benchmarks show that, while retaining only about 52--65\% of the original teacher--student discrepancy as the optimization signal, Cal-OPD consistently outperforms standard OPD and its variants across model scales.

↑