我们真的需要KL散度用于大型语言模型的同策略蒸馏吗?
Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?
浏览论文内容
中文总结 AI 辅助
本研究发现同策略蒸馏中KL散度并非必需,仅需保持朝向教师模型的更新方向,并聚焦于师生分歧大的token,据此提出C-MOPD方法,在数学和代码基准上优于MOPD。
中文摘要 AI 辅助
自从知识蒸馏问世以来,KL散度一直是蒸馏中的标准损失函数。近期,同策略蒸馏(OPD)作为一种高效的大型语言模型(LLM)后训练范式崭露头角。作为一种蒸馏方法,OPD自然继承了KL散度作为其标准损失。然而,在本工作中,我们发现KL散度可能并非OPD所必需。我们证明,仅保留更新方向就足以实现有效的OPD。只要更新方向朝向教师模型,OPD就能正常工作。更精确地说,起作用的并非每个token的方向,而是教师与学生分歧较大的小部分token的方向。我们首先证明,仅对教师概率高于学生概率的token赋予奖励(+1),对教师概率较低者赋予(-1),这仅鼓励更新朝向教师,却能几乎复现与使用反向KL的OPD相同的训练模式。我们进一步证明,只有教师-学生分歧大的小部分token的方向才是关键的,只要这些token的更新方向朝向教师,即使其他token被拉离教师,训练也能正常进行。作为这些发现的应用,我们引入了共识多教师同策略蒸馏(C-MOPD)以改进多教师同策略蒸馏(MOPD)。与MOPD将每个样本路由到单一教师并可能导致跨领域能力冲突不同,C-MOPD让每个样本都接受所有教师的监督。实验表明,C-MOPD在数学和代码基准测试上均持续优于MOPD。我们的代码可在该https URL获取。
英文摘要
Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in this work, we find that KL divergence may not be necessary for OPD. We show that simply preserving the update direction is sufficient for effective OPD. As long as the update direction is toward the teacher, OPD works. More precisely, it is not the direction of every token, but the direction of a small subset of tokens where the teacher and student disagree strongly. We first show that simply assigning a reward of (+1) to tokens where the teacher probability is higher than the student probability and (-1) where it is lower, which merely encourages updates toward the teacher, reproduces almost the same training mode as OPD with reverse KL. We further show that only the direction of a small subset of tokens with large teacher-student disagreement is critical, and training works as long as their update direction is toward the teacher, even if other tokens are pulled away from the teacher. And as an application of these findings, we introduce Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to improve Multi-Teacher On-Policy Distillation (MOPD). Unlike MOPD, which routes each sample to a single teacher and may cause capability conflicts across domains, C-MOPD lets every sample be supervised by all teachers. Experiments show that C-MOPD consistently outperforms MOPD on both math and code benchmarks. Our code is available at https://github.com/LeapLabTHU/KL-Free-OPD.
发表机构
- LeapLab, Tsinghua University(清华大学LeapLab)
- Qiuzhen College, Tsinghua University(清华大学求真书院)
- Beihang University(北京航空航天大学)
- National University of Singapore(新加坡国立大学)
- The Chinese University of Hong Kong(香港中文大学)
- E Fund Management Co., Ltd.(易方达基金管理有限公司)
机构由 AI 辅助整理,请以论文原文为准。