arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WASD:基于Wasserstein距离的大语言模型知识蒸馏

WASD: Wasserstein-based Knowledge Distillation for Large Language Models

Byeonghu Na, Donghyeok Shin, Yeongmin Kim, Mina Kang, Il-Chul Moon

arXiv 2610.07706首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型知识蒸馏忽略词元语义的问题,提出基于Wasserstein距离的WASD方法,利用词元嵌入构建代价矩阵,采用Sinkhorn散度高效优化,在指令遵循、数学推理和代码生成等任务上持续提升蒸馏性能。

AI 中文摘要

自回归大语言模型(LLMs)的能力迅速提升,但其规模增大带来了推理时大量的计算和内存开销。知识蒸馏(KD)通过将离散概率分布对齐,把知识从大型教师模型转移到较小的学生模型,提供了一种实用的解决方案。然而,现有的LLM知识蒸馏方法主要依赖基于每个词汇索引处的概率值来评估差异的散度,没有显式利用词元级别的语义信息。我们提出用于LLM的基于Wasserstein距离的知识蒸馏(WASD),该方法通过基于Wasserstein的距离,利用由词元嵌入导出的代价矩阵,纳入词元级别的语义信息。为确保计算可行性,我们采用Sinkhorn散度,并推导出一个梯度等价的优化目标,无需引入额外网络即可高效优化。跨多个LLM系列和规模的实验表明,WASD在多样任务上持续提升蒸馏性能,包括指令遵循、数学推理和代码生成。我们的结果突显了词元空间中编码的语义信息对于LLM蒸馏中有效分布对齐的重要性。实现已公开于该https URL。

英文摘要

Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate discrepancies through probability values at each vocabulary index, without explicitly leveraging token-level semantic information. We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings. To ensure computational tractability, we adopt the Sinkhorn divergence and derive a gradient-equivalent objective that can be efficiently optimized without introducing additional networks. Experiments across multiple LLM families and scales show that WASD consistently improves distillation performance on diverse tasks, including instruction following, mathematical reasoning, and code generation. Our results highlight the importance of semantic information encoded in the token space for effective distribution alignment in LLM distillation. The implementation is publicly available at https://github.com/aailab-kaist/WASD .

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑