arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GAN-Diff:耦合预训练WGAN-GP特征与条件扩散U-Net

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

Saif Ahmed, Asadullah Hil Galib, S. M. Riaz Rahman Antu, Ahmed Faizul Haque Dhrubo, Souvik Pramanik, Mohammad Abdul Qayum, Mohsin Sajjad, Mohammad Ashrafuzzaman Khan

arXiv 2608.22272首次发表:更新:

发表机构

North South University(北南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出GAN-Diff框架,将预训练WGAN-GP特征作为先验耦合到条件扩散U-Net,在CelebA数据集的高斯去噪和2倍超分辨率任务上,分别提升PSNR达4.40 dB和3.70 dB,实现稳定高效的图像复原。

AI 中文摘要

生成对抗网络(GAN)可实现高效图像生成,而扩散模型能提供高质量图像复原,但需迭代采样。本文提出一种混合GAN引导的扩散框架,使用预训练带梯度惩罚的Wasserstein GAN(WGAN-GP)作为条件扩散图像复原的特征先验。冻结的WGAN-GP生成器的中间特征通过交叉注意力被整合到扩散U-Net中,且在DDIM采样过程中保持固定。该框架在两个复原任务(高斯去噪和2倍超分辨率)上使用CelebA人脸图像进行评估。开发过程中,识别并解决了多种不稳定性来源,包括对抗学习率失衡、不当的扩散初始化、过度损坏以及参数平均不足。最终框架持续提升退化图像和低分辨率图像的质量,与各自输入基线相比,去噪性能的峰值信噪比(PSNR)提升了4.40 dB,超分辨率性能提升了3.70 dB。这些结果表明,冻结的GAN特征先验具有引导扩散模型实现稳定且有效图像复原的潜力。

英文摘要

Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but require iterative sampling. This paper presents a hybrid GAN-guided diffusion framework that uses a pretrained Wasserstein GAN with gradient penalty (WGAN-GP) as a feature prior for conditional diffusion-based image restoration. Intermediate features from the frozen WGAN-GP generator are incorporated into a diffusion U-Net through cross-attention and remain fixed during the DDIM sampling process. The framework is evaluated on two restoration tasks, Gaussian denoising and 2Xsuper-resolution, using CelebA face images. During development, several sources of instability were identified and addressed, including adversarial learning-rate imbalance, inappropriate diffusion initialization, excessive corruption, and insufficient parameter averaging. The resulting framework consistently improves the quality of both degraded and low-resolution images. In particular, it improves denoising performance by 4.40 dB in PSNR and super-resolution performance by 3.70 dB over their respective input baselines. These results demonstrate the potential of a frozen GAN feature prior to guide diffusion models toward stable and effective image restoration.

Comments7 pages, 10 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑