arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36734cs.CL

蒸馏什么重要:面向大语言模型的置信度感知选择性蒸馏

Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models

Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型知识蒸馏中教师预测不可靠的问题,提出置信度门控的CaRE-KD框架,通过自适应KL散度和认知拒绝机制提升蒸馏效果,在多项基准上取得一致增益。

中文摘要 AI 辅助

知识蒸馏(KD)通过匹配输出分布来训练容量较小的学生模型以模仿容量较大的教师模型,其隐含假设是教师模型是可靠的标准答案。在大语言模型(LLMs)中,这一假设往往不成立:教师模型的预测可能表现出高熵和幻觉,导致标准KD削弱了原本校准良好的学生模型先验。我们提出CaRE-KD,一种置信度门控的蒸馏框架,用不确定性自适应优化替代静态目标。CaRE-KD包含两个组件:一个词元级损失(CaRE-Divergence),根据教师与学生模型的置信度在前向和反向KL散度之间自适应切换;以及一个批次级认知拒绝机制(Revival),当教师模型比学生模型更不确定时抑制更新。我们提供了梯度级分析,表明这种双粒度设计引入了一种条件校准机制,这是先前的静态散度无法复现的。在实验中,跨越八个教师-学生模型对和涵盖指令遵循、对话对齐、代码生成和数学推理的十一个基准,CaRE-KD在强基线(Skewed-KL,$\alpha$-$\beta$散度)上取得了持续改进。亮点包括在指令遵循任务上平均ROUGE-L最高提升$+3.2$,在MBPP上pass@1提升$+2.1$,在GSM8k上准确率提升$+1.7$,在CollegeMath上准确率提升$+1.8$,且在LLM-as-a-judge事实性评估中持续提升(每任务最高提升$+2.5$,优于Skewed-RKL)。Revival还作为一种原则性的、与损失无关的即插即用模块,通过过滤认知上不可靠的教师监督,系统地增强了现有的蒸馏目标。

英文摘要

Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL divergence based on teacher--student confidence, and a batch-level epistemic rejection mechanism (Revival) that suppresses updates when the teacher is more uncertain than the student. We provide a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce. Empirically, across eight teacher--student pairs and eleven benchmarks spanning instruction following, chat alignment, code generation, and mathematical reasoning, CaRE-KD delivers consistent gains over strong baselines (Skewed-KL, $α$--$β$ divergence). Highlights include up to $+3.2$ average ROUGE-L on instruction-following tasks, $+2.1$ pass@1 on MBPP, $+1.7$ accuracy on GSM8k, and $+1.8$ accuracy on CollegeMath over the strongest baseline, with consistent gains in LLM-as-a-judge factuality (up to $+2.5$ per task over Skewed-RKL). Revival further acts as a principled, loss-agnostic plug-in that systematically strengthens existing distillation objectives by filtering epistemically unreliable teacher supervision.

发表机构

  • Indian Institute of Technology Delhi(印度理工学院德里分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑