更少均匀的离散扩散更强大且可扩展
Less Uniform Discrete Diffusion is More Powerful and Scalable
浏览论文内容
中文总结 AI 辅助
本文提出更少均匀扩散(LUDI)框架,通过改进训练目标和引入令牌级时间嵌入,解决了均匀扩散语言模型扩展难的问题,实现了少步生成和复杂推理,并展示了性能提升。
中文摘要 AI 辅助
尽管均匀扩散语言模型(UDLMs)代表了一种有前景的扩散范式,但对其进行扩展仍然具有挑战性。我们识别出核心障碍在于过均匀的训练目标以及采样过程中的条件-目标混淆。为解决这些问题,我们提出了更少均匀扩散(LUDI),一种新颖的UDLM框架。具体而言,我们(i)引入了一个更少均匀的损失函数,引导每个反向转移朝向干净令牌,以及(ii)为模型配备每令牌时间嵌入,提供令牌级别的损坏提示,从而实现基于置信度的少步采样。跨尺度的实验表明,LUDI产生了更清晰的监督并改善了少步生成。我们进一步将一个70亿参数的自回归模型继续训练为LUDI-7B,得到一个能够进行复杂推理的UDLM。它在每一步生成3个令牌的速度上比自回归解码有所提升,并且与掩码扩散基线相比具有竞争力的性能,这表明UDLMs在复杂生成方面的全部潜力仍有待挖掘。
英文摘要
Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM framework. Specifically, we (i) introduce a less uniform loss that directs each reverse transition toward the clean token, and (ii) equip the model with per-token time embeddings that supply token-level corruption hints, enabling confidence-based few-step sampling. Experiments across scales show that LUDI yields cleaner supervision and improves few-step generation. We further continue-train a 7B autoregressive model into LUDI-7B, resulting in a UDLM capable of complex reasoning. It achieves a 3-token-per-step speedup over AR decoding and competitive performance compared with masked diffusion baselines, revealing that the full potential of UDLMs for complex generation remains to be unlocked.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
- Huawei Foundation Model Department(华为基础模型部门)
机构由 AI 辅助整理,请以论文原文为准。