arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种通过交互解释大语言模型知识蒸馏的统一方法

A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions

Qingzhuo Wang, Ruiyang Qin, Zhenxin Qin, Wen Shen, Zhihua Wei

arXiv 2607.08776首次发表:更新:

AI 中文总结

研究大语言模型知识蒸馏有效性背后机制,提出统一方法,通过交互分解输出分数,发现共同机制是交互稀疏化,不同方法性能差异源于处理复杂交互能力,进而提出CIP损失函数,实验证明其能提升多种KD方法性能。

AI 中文摘要

尽管知识蒸馏(KD)在大语言模型(LLMs)中取得了成功,但其有效性背后的潜在机制仍不清楚。本文提出一种统一方法,利用交互探索各种KD方法的共同机制。具体而言,将LLM的输出分数分解为众多交互的总和,每个交互代表涉及一组输入变量(如单词)的非线性关系。基于分解的交互发现,各种KD方法的共同机制是交互的稀疏化,即学生模型在推理时保留较少交互,同时抑制其他交互至零效应。还发现不同KD方法的性能差异源于处理复杂交互的能力。若能使学生模型实现更高的复杂交互稀疏性,KD方法通常性能更好。基于这些见解,提出即插即用损失函数复杂交互惩罚(CIP),以在蒸馏过程中明确强化复杂交互的稀疏性。大量实验表明,集成CIP持续提高了不同KD方法在域内和分布外基准上的性能。

英文摘要

Despite the success of knowledge distillation (KD) in Large Language Models (LLMs), the underlying mechanism behind its efficacy remains unclear. In this paper, we propose a unified approach to explore the common mechanism of various KD methods using interactions. Specifically, we decompose the output score of the LLM into the sum of numerous interactions. Each interaction represents a nonlinear relationship involving a set of input variables (e.g., words). Based on the decomposed interactions, we discover that the common mechanism underlying various KD methods is the sparsification of interactions, i.e., student models retain fewer interactions for inference while suppressing other interactions to zero effects. Furthermore, we discover that the performance variance across different KD methods arises from their capabilities in handling complex interactions. A KD method typically yields better performance if it enables the student model to achieve higher sparsity of complex interactions. Motivated by these insights, we propose a plug-and-play loss function called Complex Interaction Penalty (CIP) to explicitly enforce the sparsity of complex interactions during the distillation process. Extensive experiments demonstrate that integrating CIP consistently improves the performance of diverse KD methods on both in-domain and out-of-distribution benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑