用于少步生成的隐核离散流图
Latent-Kernel Discrete Flow Maps for Few-Step Generation
浏览论文内容
中文总结 AI 辅助
提出隐核离散流图(LKF)模型,通过共享隐变量绑定分解组件实现少步生成,在两个文本基准上较似然基线提升困惑度2.1-3.3倍,M=8时优于现有少步采样器
中文摘要 AI 辅助
离散扩散模型和流匹配模型会经过多步对序列去噪,但为了保证每一步的计算成本较低,它们会将跨位置的转移分解,并独立决定每个token。当目标需要关联两个位置(如必须一致的主语和动词)时,这种独立更新会导致文本的少步生成变得极具挑战性——独立更新会分别处理这两个位置,进而需要大量函数评估来修复不匹配问题。现有的少步方法通过蒸馏或校正慢速教师模型来弥补丢失的相关性,因此继承了教师模型的质量上限。我们转而研究模型能否天然表达相关的步骤,并提出了隐核离散流图(Latent-Kernel Discrete Flow Maps, LKF),这是一种从零开始构建的流图核,是由单个共享隐变量绑定的M个分解组件的混合体。在隐变量的条件下,每个组件的计算成本较低,且对于较小的M值,可通过闭式求和对隐变量上的混合体进行计算。我们证明,单步生成会将质量分配到相关的补全结果上,且采样时间复杂度与分解模型相同,因为每个序列仅抽取一个隐变量并在整个去噪轨迹中重复使用。我们还证明,掩码扩散语言模型(Masked Diffusion Language Model, MDLM)是我们的LKF模型在M=1时的特殊情况。在One-Billion-Word(LM1B)和WikiText-103基准上进行的无条件文本生成实验表明,我们的LKF模型能学习到高度异质的组件,且生成困惑度较似然基线提升了2.1倍至3.3倍,同时未损失多样性。该增益随M值增大而提升,当M=8时,LKF模型超越了蒸馏和校正类少步采样器。源代码可在以下网址获取:this https URL
英文摘要
Discrete diffusion and flow-matching models denoise a sequence over many steps, but to keep each step cheap, they factorize the transition across positions and decide every token independently. This makes few-step generation challenging for text when the target couples two positions, such as a subject and a verb that must agree. An independent update commits to them separately, and many function evaluations are spent repairing the mismatch. Existing few-step methods buy back the lost correlation by distilling or rectifying a slow teacher, and so inherit the teacher's quality ceiling. We ask instead whether a model can express correlated steps natively, and answer with Latent-Kernel Discrete Flow Maps (LKF), a from-scratch flow-map kernel that is a mixture of M factorized components tied by a single shared latent. Conditioned on the latent, each component is cheap, and the mixture is summed over the latent in closed form for small M. We show that a single step places mass on correlated completions with the same sampling time complexity as a factorized model, since one latent is drawn per sequence and reused across the entire denoising trajectory. We also show that the Masked Diffusion Language Model (MDLM) is a special case of our LKF model at M=1. The experiments for unconditional text generation on the One-Billion-Word (LM1B) and WikiText-103 benchmarks show that our LKF model learns strongly heterogeneous components and improves generative perplexity by 2.1x to 3.3x over the likelihood baselines without losing diversity. The gain grows with M, and at M=8, it surpasses distilled and rectified few-step samplers. The source code is available at: https://github.com/mansoor181/lkf.git