均值速度匹配:重新思考扩散模型中的生成动力学
Mean Velocity Matching: Rethinking Generative Dynamics in Diffusion Models
浏览论文内容
中文总结 AI 辅助
本文提出均值速度匹配(MVM)方法,通过单一场参数化直接支持扩散模型的随机反向动力学,无需额外分数估计,并在ImageNet上达到具有竞争力的FID,统一了随机与确定性采样。
中文摘要 AI 辅助
本研究探讨了扩散模型中随机生成动力学的预测参数化问题。现有的基于速度的生成模型具有学习单一传输场的简洁性,但其标准形式是确定性的,而随机扩展通常需要额外的分数信息或中间的速度到分数的重建。为了在直接支持随机反向动力学的同时保持单场预测,本文引入了均值速度匹配(MVM)。MVM构建了一个高斯扰动过程,其中恢复导向速度$(x_0-x_t)/t$的条件期望直接构成反向SDE的漂移项。因此,单个学习场足以参数化随机反向过程,无需单独估计或重建分数。由于该速度在$t=0$附近的直接回归变得无界,MVM进一步引入了$\sqrt{t}$缩放参数化,该参数化在保持反向动力学的同时产生有界的训练目标。同一学习场还诱导出确定性概率流ODE,使得随机和确定性采样可以在统一框架内进行研究。基于Transformer的生成模型实验在ImageNet $32\times32$上以\MVMImageNetThirtyTwoNFE\\ NFE实现了\MVMImageNetThirtyTwoFID\\的FID,在ImageNet $256\times256$上以\MVMImageNetTwoFiftySixNFE\\ NFE实现了\MVMImageNetTwoFiftySixFID\\的FID。受控的SDE-ODE比较进一步表明,在极低NFE下ODE表现更好,而在函数评估充足时随机反向过程获得更低的FID。这些结果表明,MVM提供了随机反向动力学的直接单场参数化,同时保持了具有竞争力的生成质量。
英文摘要
This work studies prediction parameterization for stochastic generative dynamics in diffusion models. Existing velocity-based generative models provide the simplicity of learning a single transport field, but their standard formulation is deterministic, whereas stochastic extensions generally require additional score information or an intermediate velocity-to-score reconstruction. To retain single-field prediction while directly supporting stochastic reverse dynamics, this paper introduces Mean Velocity Matching (MVM). MVM constructs a Gaussian perturbation process for which the conditional expectation of a restoration-oriented velocity, $(x_0-x_t)/t$, directly forms the reverse-SDE drift. Consequently, a single learned field is sufficient to parameterize the stochastic reverse process without separately estimating or reconstructing the score. Because direct regression of this velocity becomes unbounded near $t=0$, MVM further introduces a $\sqrt{t}$-scaled parameterization that preserves the reverse dynamics while yielding a bounded training target. The same learned field also induces a deterministic probability-flow ODE, enabling stochastic and deterministic sampling to be studied within a unified formulation. Experiments with Transformer-based generative models achieve an FID of $\MVMImageNetThirtyTwoFID$ at \MVMImageNetThirtyTwoNFE\ NFE on ImageNet $32\times32$ and $\MVMImageNetTwoFiftySixFID$ at \MVMImageNetTwoFiftySixNFE\ NFE on ImageNet $256\times256$. Controlled SDE--ODE comparisons further show that the ODE performs better under very low NFE, whereas the stochastic reverse process achieves lower FID when sufficient function evaluations are available. These results demonstrate that MVM provides a direct single-field parameterization of stochastic reverse dynamics while maintaining competitive generation quality.
发表机构
- Chengdu University of Technology(成都理工大学)
- Peking University(北京大学)
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。