arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

失配很重要:超越token一致性的在线蒸馏

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang, Xu Wang, Difan Zou

arXiv 2608.09836首次发表:更新:

发表机构

The University of Hong Kong; University of Science and Technology of China; The Chinese University of Hong Kong(香港大学; 中国科学技术大学; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对在线蒸馏存在的退化一致性失效问题,提出TIDE方法修正师生失配,在数学推理基准上显著提升性能、缩短响应长度并减少格式错误。

AI 中文摘要

在线蒸馏(OPD)已成为现代大语言模型(LLM)后训练流程的核心组成部分,但本文揭示了其存在的一种失效模式:退化一致性,即学生模型利用重复循环实现与教师模型近乎完美的token一致性,却给出全局存在缺陷的响应。因此,本文将研究重点从一致性转向师生失配,发现失配token主要可分为两类:学生过剩token与学生不足token。学生过剩token由学生模型生成,但教师模型为其分配的概率接近0;它们的对数比修正值会无限制增长,导致更新不稳定。相反,学生不足token是教师模型偏好的,但学生模型很少采样到;它们的缺失阻碍了教师模型推理模式的迁移。为解决这些失配问题,本文提出TIDE(Token级独立不足-过剩修正),该方法应用有界的Hellinger塑形来抑制最严重的采样过剩,并采用解析的教师Top-K注入来恢复不足的概率质量,无需采样到不足token。在包含多组Qwen3师生对的数学推理基准测试中,TIDE始终优于标准OPD及近期的token选择、奖励塑形基线。此外,在师生失配严重的场景下,TIDE的提升更为显著:它将Avg@8从6.9%提高至20.3%,将平均响应长度缩短3.6倍,并大幅减少格式错误。代码可在指定URL获取。

英文摘要

On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. We therefore shift our focus from agreement to teacher-student mismatch, and find that mismatch tokens can be mainly categorized into two types: student-excess tokens and student-deficit tokens. Student-excess tokens are generated by the student but assigned near-zero probability by the teacher; their log-ratio corrections grow unbounded and destabilize the update. Student-deficit tokens, in contrast, are preferred by the teacher but rarely sampled by the student; their absence blocks the transfer of the teacher's reasoning patterns. To tackle these mismatch directions, we propose TIDE (Token-level Independent Deficit-Excess correction), which applies bounded Hellinger shaping to suppress the most severe sampled excesses and an analytic teacher top-$K$ injection to restore deficient probability mass without requiring deficit tokens to be sampled. Across mathematical reasoning benchmarks with multiple Qwen3 teacher-student pairs, TIDE consistently outperforms standard OPD and recent token-selection and reward-shaping baselines. Moreover, the gains of TIDE are more pronounced under strong teacher-student mismatch, where it improves Avg@8 from 6.9% to 20.3%, reduces average response length by a factor of 3.6, and substantially reduces formatting failures. Code is available at https://github.com/yzc-666/TIDE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑