arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知识蒸馏的统计学视角:基础、经典方法与大型语言模型扩展

A Statistical Perspective on Knowledge Distillation: Foundations, Classical Methods, and Large Language Model Extensions

Luyang Fang, Haoran Lu, Jiazhang Cai, Tao Wang, Huimin Cheng, Wenxuan Zhong, Ping Ma

arXiv 2609.33727首次发表:更新:

发表机构

University of Georgia; Boston University(佐治亚大学; 波士顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述从贝叶斯统计视角统一知识蒸馏,将教师预测视为先验信息,连接经典方法与LLM扩展,为领域提供概念框架并指出开放问题。

AI 中文摘要

知识蒸馏(KD)已成为将高容量模型的能力迁移至高效“学生”模型的关键范式,解决了计算成本、部署约束和隐私敏感场景中的关键挑战。尽管KD在实践中被广泛使用,但它常被视为一种工程技术,统一的统计学视角仍未充分发展。本综述通过提出一种统一的贝叶斯公式化表述来弥合这一差距,该表述将教师模型的预测视为先验信息。这为教师信息如何融入学生学习提供了原则性解释,并建立了与不确定性量化的严格联系。我们展示了这一基础视角如何调和经典蒸馏与生成式和基础模型系统中的现代扩展,表明当代发展仍根植于这些相同的统计学原理。通过将理论、新兴方法和多样化应用相结合,本综述提供了概念路线图,并指出了该领域未来的关键开放问题。

英文摘要

Knowledge Distillation (KD) has emerged as a vital paradigm for transferring the capabilities of high-capacity models to efficient ``student'' counterparts, addressing critical challenges in computational cost, deployment constraints, and privacy-sensitive settings. Although KD is widely used in practice, it is often viewed primarily as an engineering technique, with a unified statistical perspective remaining less developed. This review bridges that gap by presenting a unified Bayesian formulation of KD that formulates teacher predictions as prior information. This provides a principled interpretation of how teacher information is incorporated into student learning and establishes a rigorous connection to uncertainty quantification. We demonstrate how this foundational lens reconciles classical distillation with modern extensions in generative and foundation-model systems, showing that contemporary developments remain rooted in these same statistical principles. By synthesizing theory with emerging methodologies and diverse applications, this review provides a conceptual roadmap and identifies critical open problems for the future of the field.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑