arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

连续扩散语言模型的分布匹配蒸馏

Distribution Matching Distillation for Continuous Diffusion Language Models

Paul Le Van Kiem, Dario Shariatian, Umut Simsekli, Alain Durmus

arXiv 2609.40235首次发表:更新:

发表机构

Inria, PSL Research University; Cohere; CMAP, Ecole Polytechnique(法国国家信息与自动化研究所,巴黎文理研究大学; Cohere公司; 巴黎综合理工学院应用数学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出两种分布匹配蒸馏方法(Simplex-DMD和Reinforce-DMD),通过利用学生概率输出减少连续扩散语言模型的网络评估次数,在OpenWebText上分别实现49%和20%的困惑度降低。

AI 中文摘要

连续扩散语言模型并行生成所有词元,但高质量生成仍可能需要数百次网络评估(NFEs)。我们研究分布蒸馏如何通过利用学生模型的概率词元输出来降低这一成本。我们的统一公式将学生模型的输出参数化与由此产生的梯度估计器联系起来,并产生了两种具有相同学生架构和反向KL匹配目标的方法:Simplex-DMD使用连续词元松弛和路径梯度,而Reinforce-DMD使用分类采样和带有学习密度比的REINFORCE。我们为多步生成开发了这两种方法,并研究了与每种参数化相关的训练和采样选择。在OpenWebText上,对于1,024个词元的序列,Simplex-DMD在仅4次NFE下实现了45.6的生成困惑度,对应5.44纳特的单字熵,与匹配熵和采样预算下最强的扩散基线相比减少了49%。Reinforce-DMD在更大预算下改善了前沿,在256次NFE下达到14.9的生成困惑度,熵为5.00纳特,在相同比较协议下减少了20%。

英文摘要

Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and reverse-KL matching objective: Simplex-DMD uses continuous token relaxations and pathwise gradients, while Reinforce-DMD uses categorical sampling and REINFORCE with a learned density ratio. We develop both methods for multi-step generation and investigate the training and sampling choices associated with each parameterization. On OpenWebText, for sequences of 1,024 tokens, Simplex-DMD achieves a generative perplexity of 45.6 at a unigram entropy of 5.44 nats in just 4 NFEs, a 49% reduction relative to the strongest evaluated diffusion baseline at matched entropy and sampling budget. Reinforce-DMD improves the frontier at larger budgets, reaching a generative perplexity of 14.9 at an entropy of 5.00 nats with 256 NFEs, a 20% reduction under the same comparison protocol.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑