学生大语言模型是否会继承分布外鲁棒性?面向可靠知识迁移的不变性加权蒸馏
Do Student LLMs Inherit OOD Robustness? Invariance-Weighted Distillation for Reliable Knowledge Transfer
浏览论文内容
中文总结 AI 辅助
针对知识蒸馏中学生模型在分布外场景性能下降的问题,提出不变性加权蒸馏(IWD),通过预测不变性动态加权样本,降低虚假相关性影响,在多个基准上显著提升OOD鲁棒性。
中文摘要 AI 辅助
知识蒸馏(KD)旨在将高性能教师大语言模型(LLM)压缩为轻量级学生模型。然而,蒸馏后的学生模型在分布外(OOD)场景中常常表现出显著的性能下降,这一关键差距尚未得到充分探索。我们识别出导致OOD性能下降的两个复合机制:(1)数据虚假性:学生模型可能学习蒸馏数据集中的虚假相关性,而非真正的因果关系;(2)教师能力:标准KD对所有样本一视同仁,忽略了在给定样本上教师是由因果特征引导还是被虚假捷径误导。为应对这些挑战,我们提出不变性加权蒸馏(IWD),这是一个有理论基础的框架,通过基于多个合成环境中预测不变性估计的教师因果依赖度,动态地重新加权训练样本。IWD扰动虚假线索同时保留核心语义,为教师预测保持不变的样本分配更高的蒸馏权重,表明这些样本更多依赖因果特征而非虚假特征。我们从理论上证明,与标准均匀加权KD相比,IWD降低了学生的虚假到因果(S2C)梯度比率,推动学生朝向更不变的表征。在四个NLP基准(MNLI、SQuAD-v2、CoNLL-2003 NER和SST-2)上、跨两个模型家族(DeBERTa-v3和Qwen-2.5)的实验表明,IWD在OOD评估上持续优于强KD基线,同时保持具有竞争力的分布内(ID)性能。具体而言,IWD在16个OOD基准中的15个上取得最高准确率,并在NLI上将平均OOD性能较标准KD提升4.34个百分点,在QA上提升14.94个百分点。
英文摘要
Knowledge distillation (KD) aims to compress high-performance teacher LLMs into lightweight students. However, distilled students often exhibit substantial performance degradation in out-of-distribution (OOD) settings, a critical gap that remains underexplored. We identify two compounding mechanisms causing OOD performance degradation: (1) data spuriousness: students can learn spurious correlations in the distillation dataset over genuine causal relationships; and (2) teacher capability: standard KD treats all samples uniformly, ignoring whether the teacher is guided by causal features or misled by spurious shortcuts on a given sample. To address these challenges, we propose Invariance-Weighted Distillation (IWD), a theoretically grounded framework that dynamically reweights training samples using an estimate of the teacher's causal reliance derived from prediction invariance across multiple synthetic environments. IWD perturbs spurious cues while preserving core semantics, assigning higher distillation weights to samples whose teacher predictions remain invariant, indicating greater reliance on causal rather than spurious features. We theoretically show that IWD reduces the student's Spurious-to-Causal (S2C) gradient ratio compared to standard uniformly weighted KD, driving the student toward more invariant representations. Experiments on four NLP benchmarks (MNLI, SQuAD-v2, CoNLL-2003 NER, and SST-2) across two model families (DeBERTa-v3 and Qwen-2.5) demonstrate that IWD consistently outperforms strong KD baselines on OOD evaluations while maintaining competitive in-distribution (ID) performance. Specifically, IWD achieves the highest accuracy in 15 out of 16 OOD benchmarks and improves average OOD performance over standard KD by 4.34 percentage points on NLI and 14.94 percentage points on QA.
发表机构
- University of Maryland, College Park(马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。