基于表示的掩码扩散模型
Representation-based Masked Diffusion Model
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对掩码扩散模型并行采样忽略标记间依赖导致输出不连贯的问题,提出基于表示的掩码扩散模型,利用预训练编码器获取全局语义表示指导并行更新,显著提升生成质量。
AI中文摘要:
掩码扩散模型(MDMs)已成为语言建模的一种引人注目的范式,提供了高效并行文本生成的能力。然而,现有的并行采样方法通常独立地更新多个掩码标记,忽略了掩码标记之间复杂的相互依赖关系。这种独立更新机制缺乏全局协调,可能导致输出不连贯。为了解决这一局限性,我们提出了基于表示的掩码扩散模型(RMDM),该框架利用文本表示来显式编码全局语义,并帮助更精确地并行更新标记。具体来说,我们首先使用预训练编码器将文本编码到连续语义空间,并学习一个可逆变换,将表示分布归一化为高斯先验,从而在生成过程中促进高效采样。在此潜在语义表示的条件下,我们训练一个掩码扩散模型来学习条件文本分布,其中表示作为全局语义指导,协调并行标记更新并忠实逼近目标分布。实证结果表明,RMDM显著提高了生成质量,特别是在激进的少步采样场景中。
英文摘要:
Masked Diffusion Models (MDMs) have emerged as a compelling paradigm for language modeling, offering the capability for efficient parallel text generation. However, existing parallel sampling methods typically update multiple masked tokens independently and ignore the complex mutual dependencies among the masked tokens. This independent updating mechanism lacks global coordination and might lead to incoherent outputs. To address this limitation, we propose Representation-based Masked Diffusion Model (RMDM), a framework that leverages the text representation to explicitly encode global semantics and help to parallel update tokens more precisely. Specifically, we first encode text into a continuous semantic space using a pretrained encoder and learn an invertible transformation that normalizes the representation distribution to a Gaussian prior, facilitating efficient sampling during generation. Conditioned on this latent semantic representation, we train a masked diffusion model to learn the conditional text distribution, where the representation serves as global semantic guidance to coordinate parallel token updates and faithfully approximate the target distribution. Empirical results demonstrate that RMDM significantly improves generation quality, particularly in aggressive few-step sampling regimes.