arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33200cs.LGcs.CL

教自己看向何处:用于推理的在线策略注意力自蒸馏

Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin, Md Mofijul Islam

首次发表
浏览论文内容

中文总结 AI 辅助

提出在线策略注意力自蒸馏(OPASD),通过解决方案条件注意力蒸馏补充令牌级监督,在竞赛级数学基准上提升准确率、减少计算开销并加速训练。

中文摘要 AI 辅助

在线策略自蒸馏利用来自具有已验证解决方案访问权限的特权教师的密集令牌分布指导,在模型自身的轨迹上训练推理模型。这种监督转移了教师预测的内容,但没有直接转移其在先前上下文中关注的位置。我们提出了在线策略注意力自蒸馏(OPASD),它通过解决方案条件注意力蒸馏补充了令牌级监督。由于特权教师可以关注学生无法获得的已验证解决方案令牌,OPASD将教师注意力投影到学生可见的位置,并在对齐前重新归一化所得分布。在三种模型规模和四个竞赛级数学基准上,OPASD始终优于仅令牌的OPSD,平均准确率提高了4.98到8.40个百分点。OPASD还避免了仅令牌蒸馏中观察到的响应长度膨胀和性能下降,将生成的滚动令牌减少了73.9%,估计模型计算量减少了72.6%,同时训练速度提高了1.53倍。这些结果表明,解决方案条件注意力提供了一种互补的监督信号,使在线策略自蒸馏更加准确、稳定和计算高效。

英文摘要

On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.

↑