发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过扩展嵌入模型并蒸馏软标签,提升了连续扩散语言模型的可扩散性,使中等规模模型在生成困惑度上超越GPT-2-M。
AI 中文摘要
扩散语言模型(DLMs)为自回归(AR)语言生成提供了一种有前景的替代方案。近期连续扩散语言模型的进展,即将潜在扩散应用于连续文本嵌入,引发了一个实际问题:哪种嵌入能构成最佳的潜在空间,即最具可扩散性?为回答此问题,我们搜索了不同的嵌入,发现在同一系列中(从T5到T5Gemma-1再到T5Gemma-2)将嵌入模型扩展到更强的版本能大幅提升生成性能。但原始的T5Gemma-2嵌入仍非最优。它们具有极强的判别性,以至于即使是合理的替代词的嵌入也被分离,这使得生成过程易受不完美采样的影响。因此,连续扩散常常无法到达任何此类嵌入,最终落在无效嵌入上。为解决此问题,我们将T5Gemma-2蒸馏到一个学生编码器中,该编码器学习教师解码后的概率作为软标签。从这类软标签中学习使得学生将替代嵌入拉近,同时保持编码-解码机制。蒸馏后的嵌入形成了更连通且更具可扩散性的潜在空间,优于原始的T5Gemma-2嵌入。最终,我们的中等规模扩散语言模型在OpenWebText上的真实文本熵条件下达到了生成困惑度(Gen. PPL)17.8(对比真实文本困惑度15.4),在Gen. PPL上超越了GPT-2-M。
英文摘要
Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.
Comments28 pages, 12 figures. Code is available at https://github.com/la0ka1/diffusing-scaled-text-embeddings