超越空间域监督:用于多模态图像融合的关系约束空间
Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion
AI总结:
针对多模态图像融合缺乏真实标签导致监督失配的问题,提出关系约束监督范式,利用冻结预训练模型与可学习适配器构建关系空间,通过三个关系损失与对比排序优化,显著提升融合性能。
AI中文摘要:
多模态图像融合(MMIF)旨在通过整合共享信息、保留互补线索并协调跨模态冲突,将多模态信息形成单一图像。然而,由于缺乏真实融合图像,现有的MMIF监督通常使用空间域源图像或梯度变体作为替代真实值,这使得监督机制与MMIF的目标本质上不一致,并导致像素级妥协或模态偏差。为解决这一问题,我们提出了一种关系约束监督范式,将融合监督从空间域转移到学习得到的关系空间。我们不仅依赖直接的源图像近似,还进一步利用冻结的预训练表示模型作为信息提供者,并设计了一个可学习的特征适配器,将异构的DINO和CLIP特征对齐到统一的监督空间中。该适配器推断三个关系参数,即共享性、主导性和协调半径,它们定义了与MMIF目标相对应的三个损失。为了使该空间可靠,我们设计了一个针对适配器的自监督对比排序目标,并通过交替优化将其与融合网络耦合。大量实验表明,无论融合网络采用哪种主流骨干网络,所提出的监督空间都能带来显著的性能提升,提供了一种与MMIF目标更一致的监督范式。代码:此HTTP URL。
英文摘要:
Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-modal conflicts across modalities. However, due to the absence of ground-truth fused images, existing MMIF supervision commonly uses spatial-domain sources or gradient variants as surrogate ground truth, making the supervision mechanism inherently misaligned with the goal of MMIF and causing pixel-level compromise or modality bias. To address this, we propose a relation-constrained supervision paradigm that moves fusion supervision from the spatial domain to a learned relation space. Rather than relying solely on direct source approximation, we further leverage frozen pretrained representation models as information providers and design a learnable feature adapter to align heterogeneous DINO and CLIP features into a unified supervision space. The adapter infers three relation parameters, namely sharedness, dominance, and coordination radius, which define three losses corresponding to the MMIF's goal. To make this space reliable, we devise a self-supervised contrastive ranking objective tailored to the adapter and couple it with the fusion network through alternating optimization. Extensive experiments show that the proposed supervision space yields significant gains regardless of which mainstream backbone the fusion network adopts, offering a supervision paradigm better aligned with the goal of MMIF. Code: github.com/GMY628/RCS-Fusion.