arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

离散平均生成器加速扩散语言模型

Acceleration of Diffusion Language Model through Discrete Average Generator

Yidong Ouyang, Zhengyan Wan, Themis Haris, Tian Tan, Liqian Peng, Henry Li, Ziqian Lin, Jianhang Chen, Maryam Karimzadehgan, Alec Go, George Michailidis

arXiv 2609.38364首次发表:更新:

发表机构

University of California, Los Angeles; Google(加州大学洛杉矶分校; 谷歌)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出离散平均生成器,将MeanFlow扩展到连续时间马尔可夫链,通过自洽恒等式实现高效训练,在Potts模型和OpenWebText上显著加速并降低困惑度。

AI 中文摘要

离散扩散模型和流匹配已成为离散状态空间生成建模的强大框架,然而高效的少步生成仍是一个基本挑战。在这项工作中,我们引入了离散平均生成器,这是MeanFlow到连续时间马尔可夫链(CTMCs)的一个原则性扩展。类似于MeanFlow在连续空间中定义时间间隔上的平均速度场,我们将平均生成器定义为时间间隔上转移核的归一化增量。我们证明该平均生成器满足自洽恒等式,这为我们的训练目标提供了基础。我们进一步开发了与扩散语言模型标准训练范式一致的训练策略,同时保持所得目标的可处理性。当投影到每坐标边际分布时,自洽恒等式具有闭式表达式,从而实现高效的训练和推理。在Potts模型模拟中,我们的目标将$K$步采样器的总变差距离最多减少了67%。在OpenWebText上,我们的方法在8到64个采样步骤中实现了评估方法中最低的生成困惑度,同时实现了$16\ imes$加速,并在ImageNet上取得了与现有方法相当的性能。

英文摘要

Discrete diffusion models and flow matching have emerged as powerful frameworks for generative modeling over discrete state spaces, yet efficient few-step generation remains a fundamental challenge. In this work, we introduce the Discrete Average Generator, a principled extension of MeanFlow to Continuous-Time Markov Chains (CTMCs). Analogously to how MeanFlow defines an average velocity field over a time interval in continuous spaces, we define an average generator as the normalized increment of the transition kernel over a time interval. We show that this average generator satisfies a self-consistency identity, which provides the foundation for our training objective. We further develop training strategies that align with the standard training paradigm of diffusion language models while keeping the resulting objective tractable. When projected onto per-coordinate marginals, the self-consistency identity admits a closed-form expression, enabling efficient training and inference. In Potts model simulations, our objective reduces the total variation distance of the $K$-step sampler by up to 67%. On OpenWebText, our method achieves the lowest generative perplexity among the evaluated methods for 8 to 64 sampling steps while enabling a $16\times$ acceleration, and achieves comparable performance to existing methods on ImageNet.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑