arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35553cs.LGstat.ML

单纯形扩散模型

Simplex Diffusion Models

  • Google DeepMind(谷歌DeepMind)
  • EPFL(瑞士洛桑联邦理工学院)
  • UCL Gatsby(伦敦大学学院盖茨比计算神经科学中心)

机构由 AI 辅助整理,请以论文原文为准。

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

中文总结 AI 辅助

本文提出单纯形扩散模型(SDMs),通过将扩散过程提升至概率单纯形以保留类别信念,缓解信息坍缩,在文本和代码生成上优于现有离散扩散基线,且支持高效蒸馏采样。

中文摘要 AI 辅助

扩散模型通过信念状态的逐步细化,彻底改变了连续数据的生成建模。然而,这种迭代细化尚未应用于离散扩散模型,后者在中间步骤通过类别采样(信息坍缩)丢弃了不确定性。我们提出了单纯形扩散模型(SDMs),该框架将扩散过程提升到概率单纯形上,以表示对类别的信念。SDMs允许具有闭式反向转移的概率路径,并可通过简单的交叉熵损失进行训练。与早期方案(如需要求解常微分方程的Dirichlet流匹配)不同,我们引入了一种具有可调随机性水平的类DDIM采样器。由于SDMs在单纯形上的样本上操作,它们可以在去噪步骤间携带不确定性,从而缓解信息坍缩。在OpenWebText上,SDMs与强离散扩散基线竞争,在64个采样步骤中以5.46的单字熵达到17.0的GenPPL,接近真实验证数据。即使没有自条件(SC),SDMs在代码生成(TinyGSM,T=0.1;49.0%对45.8%)上也优于掩码和均匀扩散(带SC或预测-校正采样)。蒸馏至8步后,SDMs解决了GSM8K问题的32.1%,超过了128步的蒸馏离散扩散模型(21.4%)。

英文摘要

Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information collapse). We propose Simplex Diffusion Models (SDMs), a framework that lifts the diffusion process to the probability simplex to represent beliefs over categories. SDMs admit probability paths with closed-form reverse transitions and can be trained with a simple cross-entropy loss. Contrary to earlier proposals such as Dirichlet Flow Matching which requires integrating an ordinary differential equation, we introduce a DDIM-like sampler with a tunable level of stochasticity. Because SDMs operate on samples on the simplex, they can carry uncertainty across denoising steps, which mitigates information collapse. On OpenWebText, SDMs are competitive with strong Discrete Diffusion baselines, achieving $17.0$ GenPPL at $5.46$ unigram entropy in 64 sampling steps, close to real validation data. Even without Self-Conditioning (SC), SDMs outperform masked and uniform diffusion (with SC or predictor-corrector sampling) on code generation (TinyGSM, $T=0.1$; $49.0\%$ vs. $45.8\%$). Distilled down to 8 steps, SDMs solve $32.1\%$ of GSM8K problems, more than distilled Discrete Diffusion models with 128 steps ($21.4\%$).

↑