发表机构
ByteDance Seed; Princeton University; Stanford University; UCLA; UC Berkeley(字节跳动种子; 普林斯顿大学; 斯坦福大学; 加州大学洛杉矶分校; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对掩码扩散模型少步生成难的问题,提出多掩码扩散模型MultiMDM,通过特定前向过程使反向过程有起草能力,推导训练目标并制定蒸馏方案,经实验验证其为少步生成提供了有效基础。
AI 中文摘要
掩码扩散模型(MDMs)是很有前景的语言生成器家族,但高质量少步生成仍具挑战性。在MDMs中,所有前向轨迹都坍缩到单个完全掩码状态,缺乏用于一致性风格少步生成的终端熵。基于均匀状态扩散的近期少步替代方法避免了这种退化,但区分干净令牌和噪声比MDMs更难,损害建模质量和训练效率。本文提出多掩码扩散模型(MultiMDM),在前向过程中,每个干净令牌先被推向指定掩码,然后在掩码集上逐渐混合。反向过程通过在细化为干净令牌之前预测指定掩码而具有起草能力。推导了支持从预训练MDMs持续训练的封闭形式ELBO训练目标。此外,制定了纯离散状态一致性蒸馏方案,采用共享Gumbel耦合以减少路径熵。预训练和蒸馏实验表明,MultiMDM为有原则的少步生成提供了有效基础。
英文摘要
Masked diffusion models (MDMs) are a promising family of language generators, but achieving high-quality few-step generation remains challenging. In MDMs, all forward trajectories collapse to a single fully masked state, leaving no terminal entropy for consistency-style few-step generation. While recent few-step alternatives based on uniform-state diffusion avoid this degeneracy, it becomes harder to distinguish clean tokens from noise than MDMs, which usually harms modeling quality and training efficiency. In this work, we propose a multi-mask diffusion model (MultiMDM) that preserves the masking structure towards few-step generation. In the forward process, each clean token is first pushed towards a designated mask and then gradually mixes over the mask set. As a result, the backward process has a drafting capability by predicting a designated mask before refining to a clean token. We derive a closed-form ELBO training objective for MultiMDM that supports continual training from pretrained MDMs. In addition, we formulate a purely discrete-state consistency distillation scheme, with a shared-Gumbel coupling to reduce pathwise entropy. Experiments on pretraining and distillation show that MultiMDM provides an effective foundation for principled few-step generation.
Comments38 pages; Accepted at COLM 2026