SGPDFuse:语义引导的物理解耦通用多模态图像融合
SGPDFuse: Semantically-Guided Physics-Disentanglement General Multi-Modal Image Fusion
- Dalian University of Technology(大连理工大学)
- City University of Hong Kong(香港城市大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出SGPDFuse模型,通过语义-物理参数桥解耦场景属性与环境干扰,结合语义对齐机制实现多模态图像融合,在三类基准测试中达到最先进性能。
AI中文摘要:
多模态图像融合(MMIF)旨在将互补的传感器数据整合为单一表示,在保留场景固有真实性的同时消除环境干扰。现有多数方法依赖盲特征聚合,擅长信号积累但无法区分关键内容与物理退化因素。本文提出SGPDFuse,通过基于预训练视觉基础模型构建的语义-物理参数桥(SPPB)将输入映射为物理解耦的结构表示,利用固有-变化原理将不变场景属性与瞬态环境因素解耦。为引导该分解,引入语义对齐机制:通过余弦相似度将融合表示显式锚定到同一基础模型特征空间中的显著语义特征以保留关键目标,同时通过格拉姆矩阵正则化强制物理纹理保真度,严格消除不自然伪影。大量实验表明,SGPDFuse采用单一架构在红外-可见光、多聚焦及多曝光基准测试中均达到了最先进性能。
英文摘要:
Multimodal image fusion (MMIF) aims to integrate complementary sensor data into a single representation that preserves intrinsic scene reality while eliminating environmental interferences. Most existing approaches rely on blind feature aggregation, which excels at signal accumulation but fails to distinguish essential content from physical degradations. We propose SGPDFuse, which bridges this gap by mapping inputs into a physics-disentangled structural representation via a Semantic-Physical Parametric Bridge (SPPB) built on pretrained vision foundation models, utilizing the Intrinsic-Variation principle to decouple invariant scene attributes from transient environmental factors. To guide this decomposition, we introduce a Semantic Alignment mechanism: we explicitly anchor the fused representation to salient semantic features in the same foundation model feature space via cosine similarity to preserve critical targets, while enforcing physical texture fidelity through Gram-matrix regularization to strictly eliminate unnatural artifacts. Extensive experiments demonstrate that SGPDFuse achieves state-of-the-art performance across infrared-visible, multi-focus, and multi-exposure benchmarks using a single architecture.