arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

向深思熟虑的教师学习:用于数学推理的自适应在线策略自蒸馏

Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning

Jiacheng Du, Weiwei Xie, Tianyi Du, Shaoxiong Guo, Qibing Ren, Jiaheng Zhang

arXiv 2609.32667首次发表:更新:

AI 中文总结

提出自适应在线策略自蒸馏(AOPSD),通过推理DAG和自适应反馈缓解PI捷径,在数学推理基准上以更低训练成本显著提升Pass@8。

AI 中文摘要

在线策略自蒸馏(OPSD)利用教师提供的基于训练专用特权信息(PI)的令牌级反馈来训练仅基于问题的学生模型。因此,OPSD提供了密集的、在线策略的监督,并且无需更大的外部教师模型,但其有效性取决于PI的设计与利用方式。我们的初步诊断表明,教师效用与学生可学习性之间存在显著差距,其中一小部分高分歧令牌主导了蒸馏信号,且这些位置上的短教师续写比可迁移的纠正线索暴露了更多显式的PI泄漏,表明存在注入PI条件捷径的强烈意图。我们提出了自适应在线策略自蒸馏(AOPSD),它自适应地调整教师接收的信息以及其反馈对学习的影响强度。AOPSD将每个解答编码为推理有向无环图(DAG),根据学生不断演进的能力对问题进行排序,并仅将可负担的子图及其下一前沿作为PI暴露。对于高分歧令牌,AOPSD利用短教师续写作为探针,以鼓励有用的指导,同时缓解教师监督中PI条件捷径的影响。在HMMT25、AIME24、AIME25和BRUMo25上,AOPSD实现了72.5%的Pass@8,比OPSD高出6.7个百分点,比最强竞争基线高出4.2个百分点,同时以更低的成本减少了15个百分点的训练时间。

英文摘要

On-policy self-distillation (OPSD) trains a question-only student with token-level feedback from a teacher given training-only privileged information (PI). OPSD therefore provides dense, on-policy supervision, and is free of a larger external teacher, but its effectiveness rests on how PI is designed and utilized. Our preliminary diagnostics suggest a significant gap between teacher utility and student learnability, where a small fraction of high-disagreement tokens dominate the distillation signal, and short teacher continuations at these positions further expose more explicit PI leakage than transferable correction cues, indicating a strong intent on injecting PI-conditioned shortcuts. We propose Adaptive On-Policy Self-Distillation (AOPSD), which adapts what information the teacher receives and how strongly its feedback influences learning. AOPSD encodes each solution as a reasoning DAG, orders problems by the student's evolving capability, and reveals only the affordable subgraph and its next frontier as PI. For high-disagreement tokens, AOPSD utilizes short teacher continuations as probes to encourage useful guidance while mitigating PI-conditioned shortcuts among teacher supervisions. On HMMT25, AIME24, AIME25, and BRUMo25, AOPSD achieves 72.5% Pass@8, which is 6.7 percentage points above OPSD and 4.2 above the strongest competing baseline while reducing 15 percentage points of training time at lower cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑