arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29846cs.CLcs.LG

面向影响力引导的知识蒸馏:解决采样 token 在线策略蒸馏中的多样性瓶颈

Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation

  • University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
  • University of Science and Technology of China(中国科学技术大学)
  • Shanghai University of Finance and Economics(上海财经大学)

机构由 AI 辅助整理,请以论文原文为准。

Run Yang, Runpeng Dai, Jie Sun, Jielei Zhang, Fan Zhou, Hongtu Zhu, Peiyi Li, Longwen Gao

AI总结:

该研究针对采样 token 在线策略蒸馏的多样性瓶颈,提出 IDA-OPD 方法,通过保留熵扩张更新、替换熵收缩更新,在降低成本的同时提升 pass@k,继承教师模型的多样性,且保持普通 OPD 的 pass@1 性能。

AI中文摘要:

采样 token 在线策略蒸馏(OPD)利用学生生成的 token 高效地将能力从教师模型迁移到学生模型,仅需教师模型对采样 token 的概率即可。但它常出现多样性蒸馏失效问题:学生模型的 pass@1 指标提升,而 pass@k 指标停滞,无法继承教师模型的多样性。为解释该现象,本文提出一阶局部熵影响力,这是一种带符号的一阶代理,可将每次更新的熵效应解耦为教师-学生对数概率差距与学生局部概率结构,并通过实验将熵收缩与负影响力位置关联。受此启发,本文提出影响力引导的自适应在线策略蒸馏(IDA-OPD):不依赖计算成本高昂的全词汇前向 KL 目标,而是保留熵扩张更新,并用发散自适应优势收缩替换熵收缩更新,仅使用教师模型的采样 token 对数概率。在面向推理的蒸馏实验中,IDA-OPD 持续提升 pass@k,通过蒸馏继承教师模型的多样性,以更低成本达到最强的教师信息方法的性能,且在不使用全词汇教师信息的情况下,广泛保持了普通 OPD 的 pass@1 指标。

英文摘要:

Sampled-token on-policy distillation (OPD) efficiently transfers capabilities from teacher to student using student-generated tokens, requiring teacher probabilities only for sampled tokens. Yet it frequently suffers from diversity distillation failure: the student's pass@1 improves while its pass@$k$ plateaus, failing to inherit the teacher's diversity. To explain this, we introduce First-Order Local Entropy Influence, a signed first-order proxy that decouples each update's entropy effect into the teacher--student log-probability gap and the student's local probability structure, and empirically links entropy contraction to negative-influence positions. Motivated by this, we propose Influence-Directed Adaptive On-Policy Distillation (IDA-OPD): rather than relying on costly full-vocabulary Forward-KL objectives, it preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability. Experiments on reasoning-oriented distillation show IDA-OPD consistently improves pass@$k$, inheriting the teacher's diversity through distillation, matches the strongest teacher-informed methods at strictly lower cost, and broadly maintains vanilla OPD's pass@1, all without full-vocabulary teacher information.

↑