arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33020cs.CV

残差扩散隐式模型

Residual Diffusion Implicit Models

João Guerreiro, Pedro Tomás, Helena Aidos, Jacinto C. Nascimento

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散模型在逆问题中的初始化错位和长链幻觉问题,提出残差扩散隐式模型(RDIM),显式建模HQ-LQ残差,通过非马尔可夫隐式采样实现少步重建,并引入可控方差机制,在去噪和超分辨率基准上超越现有方法。

中文摘要 AI 辅助

扩散模型在多个任务上取得了最先进的结果。然而,在逆问题中,从纯高斯噪声的标准初始化会使生成过程与真实世界的退化错位。更近期的方法如扩散桥施加严格的端点约束,并且通常需要较长的反向过程,这容易产生幻觉。另一种一致性模型提供噪声不变的单步映射,但缺乏固有的方差建模,并且在严重损坏下可能退化。因此,提出了残差扩散隐式模型(RDIMs),构成一个广义框架,显式地建模高质量(HQ)和低质量(LQ)图像之间的残差,使前向过程与实际退化对齐。推导了一种非马尔可夫隐式反向采样器,可以跳过中间时间步,实现准确的少步甚至单步重建,同时减轻长扩散链固有的幻觉。RDIM还引入了一种可控方差机制,在确定性和随机采样之间插值,平衡保真度和多样性。此外,它使得在需要时能够直接使用感知损失。在去噪和超分辨率基准上的实验表明,RDIMs在PSNR、SSIM和LPIPS方面持续优于最先进的方法,包括桥模型和一致性模型,减少了幻觉,同时只需要少量采样步骤(通常仅一步)。结果将RDIMs定位为广泛图像恢复任务的高效解决方案。

英文摘要

Diffusion models achieve state-of-the-art results across multiple tasks. However, in inverse problems, standard initialization from pure Gaussian noise misaligns the generative process with real-world degradations. More recent methods such as diffusion bridges impose strict endpoint constraints and often require long reverse processes that are prone to hallucinations. Alternative consistency models provide noise-invariant, one-step mappings but lack inherent variance modeling and can degrade under severe corruption. Hence, residual diffusion implicit models (RDIMs) are proposed, constituting a generalized framework that explicitly models the residuals between high-quality (HQ) and low-quality (LQ) images, aligning the forward process with the actual degradation. A non-Markovian implicit reverse sampler is derived, which can skip intermediate timesteps, enabling accurate few-step or even single-step reconstruction, while mitigating the hallucinations inherent to long diffusion chains. RDIM also introduces a controllable variance mechanism that interpolates between deterministic and stochastic sampling, balancing fidelity and diversity. Furthermore, it enables the straightforward use of perceptual losses, when needed. Experiments on denoising and super-resolution benchmarks demonstrate that RDIMs consistently outperforms the state of the art, including bridge and consistency models, in terms of PSNR, SSIM, and LPIPS, reducing hallucinations while requiring only a few sampling steps (often just one). The results position RDIMs as an efficient solution for a broad range of image restoration tasks.

补充信息

↑