arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

特权解还是上下文诱导的教师行为?剖析同策略自蒸馏

Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar, Junpei Komiyama

arXiv 2608.09228首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence; Nagoya University; RIKEN AIP(穆罕默德·本·扎耶德人工智能大学; 名古屋大学; 理化学研究所人工智能项目)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究剖析同策略自蒸馏(OPSD)的性能来源,提出 OP²SD 方法,实验显示其优于基础模型、与 OPSD 相当,表明教师的上下文诱导行为是 OPSD 性能提升的重要因素。

AI 中文摘要

同策略自蒸馏(OPSD)通常被解释为特权信息的迁移:教师观察目标问题的已验证解,监督学生的轨迹。但该解释混淆了两种效应:参考解不仅揭示当前实例的答案,还改变了教师提供 token 级监督的上下文。我们用 OP²SD(跨问题同策略自蒸馏)研究特定目标特权的作用,该方法将配对参考解替换为来自不同示例的问题和解,同时保留学生 rollout、教师和蒸馏目标。在三个模型和三个数学基准上,OP²SD 优于基础模型,与 OPSD 表现相当。OP²SD 的成功表明,OPSD 的性能提升不一定来自参考解的访问,教师的上下文诱导行为是重要因素。

英文摘要

On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. However, this interpretation conflates two effects. The reference solution not only reveals the answer to the current instance but also changes the context under which the teacher provides token-level supervision. We investigate the role of target-specific privilege with $\mathrm{OP}^{2}\mathrm{SD}$ (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objective. Across three models and three mathematics benchmarks, $\mathrm{OP}^{2}\mathrm{SD}$ improves over the base model, remains competitive with OPSD. The success of $\mathrm{OP}^{2}\mathrm{SD}$ implies that OPSD gains do not necessarily come from access to the reference solution, and that the teacher's context-induced behavior is an important factor.

CommentsWorking progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑