从边际到联合的票证:扩散语言模型中一步生成块的条件噪声蒸馏
A Ticket from Marginals to Joints: Coupled-Noise Distillation for One-Step Block Generation in Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
针对扩散语言模型多步生成块的低效问题,提出CONDOR方法,通过条件噪声蒸馏实现单步生成块,在TinyStories上显著提升一步合法性。
中文摘要 AI 辅助
自回归语言模型每次前向传播只生成一个词元;扩散语言模型则通过多步生成一个词元块。我们研究是否可以在单次前向传播中生成一个块。我们通过一种噪声条件掩码去噪器来研究这一问题:向掩码嵌入添加一个与数据无关的高斯噪声场,这样原则上每个采样的噪声场会选择该块的一个联合模式。训练此类模型的既定方法是为每个示例采样多个噪声场,并通过胜者全得或重要性加权让它们竞争数据。这仅给噪声提供了粗略的控制:在我们的实验中,噪声携带的信息大致随竞争噪声场数量的对数增长,并且在测试的模型规模下,一步输出很少保持一致。我们提出了CONDOR(一步读出条件噪声蒸馏)。一个噪声条件教师模型使用随机数量的掩码位置和胜者全得进行训练。一个学生模型提出一步块,保留选定的词元,并从教师在同一噪声场下多步重填其他位置所获得的块中学习;一个基于真实数据的无噪声掩码语言模型项锚定学生模型。在TinyStories上的人工评估显示,一步合法性有大幅提升,而不同的噪声场仍能产生不同的块,每个块只需一次前向传播。
英文摘要
Can a diffusion language model generate a coherent token block in one forward pass? Masked models already predict every position at once, but each prediction is the marginal distribution given the visible context, so the tokens can be mutually inconsistent and later steps revise those already committed. We introduce CONDOR (Coupled-Noise Distillation for One-Step Readout), trained from scratch to map different noise samples to different coherent blocks. Initially, random noise is not naturally paired with a target. Winner-take-all supervision lets different samples specialize, and self-distillation trains the one-pass output to match the refined coherent block. TinyStories experiments show diverse, coherent continuations over successive blocks, one forward pass each. Qualitative MNIST experiments show that the same approach can extend to multimodal generation, such as text-to-image and unconditional text-and-image generation.
发表机构
- School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
- Zhongguancun Academy(中关村学院)
机构由 AI 辅助整理,请以论文原文为准。