发表机构
University of Cambridge; MediaTek Research; University of Warwick(剑桥大学; 联发科研究院; 华威大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出K-DCT协方差模型改进DDPM,降低计算复杂度,在少采样步数下于多数据集上实现优于现有SOTA的采样性能。
AI 中文摘要
去噪扩散概率模型(DDPMs)的采样过程可通过利用去噪后验协方差近似形式的二阶信息加速,从而在更少但更大的采样步数中生成可接受质量的样本。此前利用此类信息的尝试对协方差进行了大幅简化(如对角化处理),但未充分考虑自然图像的特殊统计结构——自然图像中像素间、颜色通道间存在强非对角相关性,且具有幂律频谱的慢衰减特性。本文提出一种新型协方差模型以捕捉这些特征:Kronecker-DCT(K-DCT)模型,采用颜色间协方差的Kronecker分解,结合离散余弦变换(DCT)在频域建模空间协方差;DCT的使用将计算复杂度从二次降至对数线性,使每步去噪的计算与内存开销可忽略。通过在CIFAR-10、Celeb-A、ImageNet及LSUN数据集上,利用预训练的评分模型学习K-DCT结构的去噪后验协方差摊销,我们的方法在FID和似然指标上均优于现有SOTA去噪采样器,尤其在少去噪步数的场景中表现突出。
英文摘要
The sampling process of Denoising Diffusion Probabilistic Models (DDPMs) can be accelerated by leveraging second-order information in the form of approximations to the denoising posterior covariance -- allowing samples of acceptable quality to be produced in fewer but larger sampling steps. Previous attempts at using such information have used drastic (e.g.\ diagonal) simplifications of the covariance. These do not do justice to the peculiar statistical structure of natural images, which exhibit strong non-diagonal correlations between pixels and color channels, and a slow-decaying power-law frequency spectrum. Here, we develop a novel covariance model that captures these features. Our Kronecker-DCT (K-DCT) model uses a Kronecker-factored decomposition of inter-color covariances and spatial covariances modeled in the frequency domain using the Discrete Cosine Transform (DCT). The use of the DCT reduces the computational complexity from quadratic to log-linear, resulting in negligible computational and memory overhead in each denoising step. By learning K-DCT-structured amortizations of the denoising posterior covariance using pre-trained score models on CIFAR-10, Celeb-A, ImageNet and LSUN datasets, we show improved performance compared to previous SOTA denoising samplers, both in terms of FID and likelihoods, especially in the regime of few denoising steps.
Comments29 pages, 11 figures
Journal refTMLR 06/2026