发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大语言模型通用知识蒸馏的空白,提出KDFP这一白盒通用知识蒸馏方法,在9个基准测试中性能优于现有方法1.6%-4.9%,训练效率最高提升99.1%。
AI 中文摘要
知识蒸馏是一种成熟技术,通过用更大、更强大的教师模型的表示训练小型高效的学生模型,来提升学生模型的能力。近期大语言模型(LLM)蒸馏的大量工作聚焦于蒸馏后训练阶段习得的能力,如指令跟随、思维链推理和工具使用,这在LLM通用知识蒸馏领域留下了巨大研究空白,而通用知识蒸馏对开发适合边缘设备部署的高效、隐私保护系统至关重要。本文采用第一性原理方法,评估过往研究的经验并开展新探索,以开发适用于现代LLM的蒸馏方法。我们提出KDFP,一种用于LLM的白盒通用知识蒸馏新方法。我们在9个基准测试中证明,KDFP的性能优于现有方法1.6%至4.9%,且通过临时参数减少将训练效率提升了高达99.1%。
英文摘要
Knowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models. Much of the recent work in the distillation of large language models (LLMs) has focused on distilling abilities learned during post-training, such as instruction following, chain-of-thought reasoning, and tool usage. This has left a large research gap in general knowledge distillation for LLMs, which is essential for developing efficient and private systems suitable for deployment on edge devices. We take a first-principles approach, evaluating previous lessons from prior works and conducting new explorations to develop a distillation methodology suitable for modern LLMs. We present KDFP, a novel methodology for white-box general knowledge distillation in LLMs. We demonstrate that KDFP outperforms existing methods by 1.6% $-$ 4.9% across 9 benchmarks while increasing training efficiency by up to 99.1% through ephemeral parameter reduction.
Comments30 pages, 6 figures, submitted to COLM 2026