arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

采样token反向KL策略上蒸馏的token级分析

A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation

Bing Shao, Jiazheng Zhang, Long Ma, Yujiong Shen, Senjie Jin, Xin Guo, Yuming Yang, Mingxu Chai, Zhiheng Xi, Boyang Liu, Junlin Shang, Tao Gui, Qi Zhang, Xuanjing Huang

arXiv 2608.25643首次发表:更新:

发表机构

College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究分析反向KL策略上蒸馏的梯度特性,提出Surprise-aware Reweighting(SuRe)加权规则,在Qwen3学生模型上提升了数学指标且域外基准无明显下降。

AI 中文摘要

策略上蒸馏(OPD)利用冻结教师模型的token级信号,在自身轨迹上监督学生模型,但采样损失如何在token间分配更新仍未被充分理解。我们分析了反向KL的每token K2估计量对学生模型logits的梯度,该梯度的L1范数可分解为教师与学生log概率差距的绝对值,以及当采样token在学生模型下的概率越低时会增大的学生侧softmax因子。在我们的数学蒸馏运行中,这些每token范数呈现高度非均匀性:低学生概率token在其总和中占比过高,且在教师-学生差距较大的token中也更为富集。作为该分析所建议的轻量级干预措施,我们研究了Surprise-aware Reweighting(SuRe,即惊喜感知重加权),这是一种分离且有界的加权规则,可进一步放大这种现有的分配方式。在两种Qwen3学生模型规模上,SuRe相比普通OPD提升了多项数学指标,且在选定的域外基准上未出现明显性能下降。因此,我们的主要贡献是对采用K2估计量训练的反向KL OPD进行了梯度级表征,SuRe是该表征的一个实证实例。

英文摘要

On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. The $\ell_1$ norm of this gradient factorizes into the absolute teacher--student log-probability gap and a student-side softmax factor that grows as the sampled token becomes less likely under the student. In our math-distillation runs, these per-token norms are highly non-uniform: low-student-probability tokens account for a disproportionate share of their sum and are also enriched in large teacher--student gaps. As a lightweight intervention suggested by this analysis, we study Surprise-aware Reweighting (SuRe), a detached, bounded weighting rule that further amplifies this existing allocation. Across two Qwen3 student scales, SuRe improves several math metrics over vanilla OPD and shows no clear degradation on the selected out-of-domain benchmarks. Our primary contribution is therefore a gradient-level characterization of reverse-KL OPD trained with the K2 estimator, with SuRe as one empirical instantiation.

Comments16 pages, 7 figures; v2 adds Boyang Liu and Junlin Shang to the author list; scientific content unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑