知识蒸馏何时会产生负面影响?低资源语言摘要的可靠性感知蒸馏
When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization
浏览论文内容
中文总结 AI 辅助
研究低资源语言摘要中知识蒸馏的问题,提出CHAD和EWAD+CPDP两种可靠性感知蒸馏方法,在BanSum基准测试中显著优于标准KD,EWAD+CPDP在多种语言上也有良好表现,还发布相关资源助力研究。
中文摘要 AI 辅助
知识蒸馏(KD)是压缩序列到序列模型的标准方法,但其样本层面的效果很少被研究。在BanSum孟加拉语摘要基准测试中,发现标准KD相比交叉熵基线仅将ROUGE-L提高了+0.0003,且约51.3%的训练样本在标准KD下会对学生验证损失产生负面影响。提出了两种互补的可靠性感知蒸馏方法。CHAD通过与验证损失方向的梯度对齐来衡量样本KD的有用性,并训练一个轻量级门来将这种反事实判断推广到整个训练集。EWAD+CPDP将令牌级熵加权自适应蒸馏与来自第二个词汇不兼容教师的容量比例几何约束相结合。在BanSum上,两种方法都显著优于标准KD:CHAD的ROUGE-L提高了+0.0173,EWAD+CPDP提高了+0.0219,而标准KD本身仅提高了+0.0003;尽管参数仅为60M,但两者都优于微调后的Qwen 2.5-3B模型(大50倍)。进一步在15种类型多样的XL-Sum语言上评估了更强的EWAD+CPDP方法,在10/15种语言上超过了仅使用交叉熵的基线;在两位教师贡献互补信号的地方收益最可靠,在他们的目标语言覆盖饱和或联合较弱的地方收益最弱。还发布了代码和训练模型以支持可重复性和对选择性蒸馏的进一步研究。
英文摘要
Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined. On the BanSum Bangla summarization benchmark, we find that standard KD improves ROUGE-L by only +0.0003 over a cross-entropy baseline, and that approximately 51.3% of training samples are estimated to actively harm student validation loss under standard KD. We propose two complementary reliability-aware distillation methods. CHAD (Counterfactual Harm-Aware Distillation) measures per-sample KD usefulness via gradient alignment with the validation loss direction and trains a lightweight gate that generalizes this counterfactual judgment to the full training set. EWAD+CPDP combines token-level entropy-weighted adaptive distillation with a capacity-proportional geometric constraint from a second, vocabulary-incompatible teacher. On BanSum, both methods substantially outperform standard KD: CHAD by +0.0173 ROUGE-L and EWAD+CPDP by +0.0219 ROUGE-L, where standard KD itself improves ROUGE-L by only +0.0003; despite using only 60M parameters, both outperform a fine-tuned Qwen 2.5-3B model (50x larger). We further evaluate the stronger method, EWAD+CPDP, across 15 typologically diverse XL-Sum languages organised into three sets, beating the CE-only baseline on 10/15 languages; gains are most reliable where the two teachers contribute complementary signal, and weakest where they have saturated or jointly weak target-language coverage. We release code and trained models to support reproducibility and further research on selective distillation.
发表机构
- BRAC University(BRAC大学)
机构由 AI 辅助整理,请以论文原文为准。