发表机构
LMU Munich; Munich Center for Machine Learning; Google DeepMind; UCL(慕尼黑大学; 慕尼黑机器学习中心; 谷歌DeepMind; 伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文改进分布扩散模型,通过延迟粒子扩展和引入随时间变化的评分规则调度,使DDM训练在ImageNet-256上实用,实现4步4.48 FID和50步2.38 FID,且FID不随采样预算退化,并迁移至文本到图像生成。
AI 中文摘要
分布扩散模型(DDMs)用通过评分规则目标训练的分布去噪器取代了标准均值预测去噪器,学习 $p(x_1 \mid x_t)$ 的随机近似,而不是其条件均值。然而,将DDM扩展到现代图像生成设置面临两个障碍:(i)多粒子训练的开销随粒子数量增加而增加;(ii)DDM使用全局固定的评分规则超参数,迫使在采样预算之间进行单一权衡。我们通过将粒子扩展推迟到Transformer的后期层来缓解这些限制,并通过引入由Biroli2024的动力学机制启发的随时间变化的评分规则调度来缓解超参数权衡。结合基于DiT的潜在设置,这些更改使DDM训练在类别条件ImageNet-$256^2$上变得实用,在4步时达到4.48 FID,在50步时达到2.38 FID,使用DiT-XL/2,从单个模型从头开始单阶段训练,无需教师、自蒸馏或JVPs。结果是一个随机少步生成器,其FID不会随着采样预算从4增加到50 NFE而退化,并且相同的配方可迁移到文本到图像生成。代码和预训练模型可在该https URL获取。
英文摘要
Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to $p(x_1 \mid x_t)$ rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two obstacles: (i) multi-particle training incurs overhead that scales with the number of particles, (ii) DDMs use globally fixed scoring rule hyperparameters, forcing a single trade-off across sampling budgets. We mitigate these limitations by deferring particle expansion to late transformer layers, and the hyperparameter trade-off by introducing time-dependent scoring rule schedules informed by the dynamical regimes of~\citet{Biroli2024}. Combined with a DiT-based latent setup, these changes make DDM training practical on class-conditional ImageNet-$256^2$, achieving 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2, from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs. The result is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Code and pre-trained models available at https://github.com/CompVis/iDDM.