arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28235cs.CV

Diff-RF:通过退化感知学习实现图像配准与融合的相互增强

Diff-RF: Mutually Reinforced Image Registration and Fusion via Degradation-Aware Learning

  • Wuhan University(武汉大学)
  • Southeast University(东南大学)

机构由 AI 辅助整理,请以论文原文为准。

Xunpeng Yi, Zaixi Du, Qinglong Yan, Yibing Zhang, Han Xu, Jiayi Ma

AI总结:

Diff-RF提出一种退化感知的相互增强扩散框架,通过模态内恢复与跨模态扩散配准融合的耦合,在复杂退化条件下实现未配准多模态图像的高质量配准与融合。

AI中文摘要:

图像配准与融合旨在从未对齐的多模态源图像中建立空间对应关系,并整合互补信息。然而,在真实成像场景中,源图像常受到复杂多样的退化影响,如低光照、噪声等,这严重阻碍了配准与融合的有效性。为解决这一问题,我们提出了一种通过退化感知学习实现的相互增强的图像配准与融合扩散框架,称为Diff-RF。该框架探索了退化条件下配准-融合与信息恢复之间的内在耦合,从而能够在复杂退化条件下实现未配准图像的高质量融合。首先,设计了模态内恢复模块,通过利用每种模态内的信息来减轻特定模态的退化,从而为配准提供更可靠的结构表示,并促进后续的跨模态融合。其次,我们开发了一个跨模态扩散配准与融合模块,在配准与融合之间建立双向交互。通过将融合导出的视觉线索和基于对应关系的几何条件整合到扩散过程中,所提出的框架逐步细化空间对齐,并利用跨模态互补信息实现协同增强。退化感知的信息恢复与配准和融合的协同优化并非被当作独立组件,而是紧密耦合,从而实现了整体性能的提升。在多个扩展数据集上的大量实验表明,Diff-RF在各种退化场景下均实现了优越的配准精度和融合质量,展现出强大的鲁棒性和泛化能力。

英文摘要:

Image registration and fusion aim to establish spatial correspondences from misaligned multi-modal source images, and integrate complementary information. However, in real-world imaging scenarios, source images are often affected by complex and diverse degradations, such as low illumination, noise, etc., which severely hinder the effectiveness of registration and fusion. To address this issue, we propose a mutually reinforced image registration and fusion diffusion framework via degradation-aware learning, termed Diff-RF. It explores the intrinsic coupling between registration-fusion and information restoration in the degradation conditions, enabling high-quality fusion of unregistered images under complex degradation conditions. First, the intra-modal restoration module is designed to alleviate modality-specific degradations by leveraging information within each modality, thereby providing more reliable structural representations for registration and facilitating subsequent cross-modal fusion. Second, we develop a cross-modal diffusion registration and fusion module that establishes bidirectional interaction between registration and fusion. By integrating fusion-derived visual cues and correspondence-based geometric conditions into the diffusion process, the proposed framework progressively refines spatial alignment and exploits cross-modal complementary information to achieve collaborative enhancement. Rather than treating them as independent components, degradation-aware information restoration and the collaborative optimization of registration and fusion are tightly coupled, achieving overall performance improvements. Extensive experiments on multiple extended datasets demonstrate that Diff-RF achieves superior registration accuracy and fusion quality under various degraded scenarios, exhibiting strong robustness and generalization ability.

↑