SelFusion:面向扩散语言模型的自蒸馏方法
SelFusion: Self-distillation for Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
本文针对扩散语言模型(DLM)生成质量差的问题,提出无需外部教师模型的自蒸馏框架SelFusion,通过双向蒸馏两种掩码模式实现性能提升,在指令跟随任务中效果优于传统知识蒸馏方法。
中文摘要 AI 辅助
扩散语言模型(DLM)缓解了自回归(AR)大型语言模型(LLM)固有的延迟瓶颈,但生成质量下降限制了其实际应用。尽管知识蒸馏(KD)是提升性能的有前景方向,但实证发现,直接应用传统KD仅能带来微小提升,甚至会降低生成质量。基于这些观察,本文提出一种面向DLM的新型自蒸馏框架SelFusion,无需外部教师模型即可实现有效KD:该框架执行两次具有不同掩码程度的前向传播,定义掩码概率更大的为硬模式,掩码概率更小的为易模式;但易模式并非总是比硬模式更准确,且可能对错误标记过于自信,因此引入两种模式间的双向KD,可基于标记级正确性动态确定蒸馏方向。在指令跟随任务上的实验结果表明,所提自蒸馏方法显著优于使用外部LLM和DLM教师的其他KD方法,在多种配置下,经SelFusion训练的学生模型甚至超越了LLM教师的性能,为提升DLM生成质量提供了实用路径。源代码可在该https URL获取。
英文摘要
Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) large language models (LLMs), but their degraded generation quality limits practical applicability. Although knowledge distillation (KD) can be a promising direction for improving performance, we empirically find that naively applying conventional KD yields only marginal gains, or even degrades generation quality. Based on these observations, we propose a novel self-distillation framework for DLMs, namely SelFusion. To enable effective KD without an external teacher model, SelFusion performs two forward passes with different masking levels, defining the hard mode with a larger masking probability and the easy mode with a smaller masking probability. However, the easy mode is not always more accurate than the hard mode and can be overconfident on incorrect tokens. Thus, we introduce bidirectional KD between the two modes, which can dynamically determine the distillation direction based on token-level correctness. Experimental results on instruction-following tasks show that the proposed self-distillation substantially outperforms other KD methods with external LLM and DLM teachers. In many configurations, the student trained with SelFusion even surpasses the performance of the LLM teacher, providing a practical path toward improving DLM generation quality. Source code can be found at https://github.com/scai-research/SelFusion_official
发表机构
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。