arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.01395cs.CL

重新思考选择性知识蒸馏

Rethinking Selective Knowledge Distillation

  • Blavatnik School of Computer Science(Blavatnik计算机科学学院)
  • Tel Aviv University(特拉维夫大学)

机构由 AI 辅助整理,请以论文原文为准。

Almog Tavor, Itay Ebenspanger, Neil Cnaan, Mor Geva

更新

AI总结:

本文提出SE-KD方法,通过学生熵引导位置选择,提升大型语言模型在准确性、下游任务适应性和内存效率上的表现,同时减少训练时间和存储需求。

AI中文摘要:

随着越来越多的努力旨在改进大型语言模型(LLMs)中的知识蒸馏(KD),替代密集教师监督的选择性蒸馏被采用,它使用token位置、词汇类别或训练样本的子集进行监督。然而,仍然不清楚哪些重要信号、选择策略及其相互作用最有效。在本工作中,我们重新审视在自回归LLMs中何处以及如何进行蒸馏。我们沿着位置、类别和样本轴解构选择性KD,并系统比较重要信号和选择策略。然后,基于此分析,我们识别出未充分利用的机会,并引入学生熵引导的位置选择(SE-KD)。在一系列基准测试中,SE-KD在准确性、下游任务适应性和内存效率方面通常优于密集蒸馏。扩展此方法到类别和样本轴(SE-KD 3X)会产生互补的效率增益,使离线教师缓存成为可能。在实践中,这减少了70%的壁时间,减少了18%的峰值内存,并在不牺牲性能的情况下将存储使用量降低了80%。

英文摘要:

Growing efforts to improve knowledge distillation (KD) in large language models (LLMs) replace dense teacher supervision with selective distillation, which uses a subset of token positions, vocabulary classes, or training samples for supervision. However, it remains unclear which importance signals, selection policies, and their interplay are most effective. In this work, we revisit where and how to distill in autoregressive LLMs. We disentangle selective KD along the position, class, and sample axes and systematically compare importance signals and selection policies. Then, guided by this analysis, we identify underexplored opportunities and introduce student-entropy-guided position selection (SE-KD). Across a suite of benchmarks, SE-KD often improves accuracy, downstream task adherence, and memory efficiency over dense distillation. Extending this approach across the class and sample axes (SE-KD 3X) yields complementary efficiency gains that make offline teacher caching feasible. In practice, this reduces wall time by 70% and peak memory by 18%, while cutting storage usage by 80% over prior methods without sacrificing performance.

↑