arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

E$^2$-OPSD:驯服在线策略自蒸馏中的熵过冲

E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation

Yifei Liu, Minghao Fang, Xinyu Gu, Chengkai Yao, Mengdi Liu, Tengfei Ma, Jiangbin Zheng, Chang Yu, Zhangyang Gao

arXiv 2610.05048首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory; Shanghai Jiaotong University; Visual Information Processing and Learning, ICT, CAS; Hunan University; Westlake University; Nanjing University; The Chinese University of Hong Kong; Zhejiang University(上海人工智能实验室; 上海交通大学; 中国科学院计算技术研究所视觉信息处理与学习; 湖南大学; 西湖大学; 南京大学; 香港中文大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线策略自蒸馏中的熵过冲问题,提出E$^2$-OPSD方法,通过示例引导教学和熵感知蒸馏提升数学推理性能,在mean@16和pass@8上分别最高提升4.9和5.5个百分点。

AI 中文摘要

在线策略自蒸馏(OPSD)无需第二个模型即可提供密集的令牌级监督:一个网络既充当教师(使用参考答案)又充当学生(仅使用问题)。我们识别出这种方法的特定失败模式。在训练过程中,学生令牌的熵超过教师并持续保持较高水平,我们将这种模式称为熵过冲。我们将其追溯到蒸馏的两个方面。以参考答案为条件的教师在其答案导向的推理路径上很自信,但这种自信难以迁移到学生生成的前缀上,使得其监督过度依赖特定答案的线索,而非可复用的推理模式;同时,OPSD使用的前向KL不断扩散学生的预测分布,却未将其拉回。我们引入E$^2$-OPSD来解决这两个原因。示例引导教学将当前答案替换为检索到的已解决相邻问题,提供可迁移的推理指导而不揭示最终答案,并更好地匹配学生可达状态。熵感知蒸馏利用学生-教师熵差来确定每个令牌修正的方向和强度。E$^2$-OPSD在数学推理上相比OPSD在mean@16上最高提升4.3个百分点,而域外评估显示相对于相应基础模型在mean@16上最高提升4.9个百分点,在pass@8上最高提升5.5个百分点。尽管有这些提升,E$^2$-OPSD仍然简单,不需要额外的前向传递或网络。

英文摘要

On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E$^2$-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E$^2$-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E$^2$-OPSD remains simple, requiring no additional forward passes or networks.

Comments24 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑