arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当积分遇见分解:一种用于多模态图像融合的信号级自监督特征分解范式

When Integral Meets Decomposition: A Signal-Level Self-Supervised Feature Decompose Paradigm for Multi-Modal Image Fusion

Zeyu Wang, Jiayu Wang, Haiyu Song, Haoran Duan

arXiv 2609.39004首次发表:更新:

发表机构

College of Computer Science and Engineering, Dalian Minzu University; School of Artificial Intelligence and Robotics, Hunan University; Department of Automation, Tsinghua University(大连民族大学计算机科学与工程学院; 湖南大学人工智能与机器人学院; 清华大学自动化系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种信号级自监督特征分解范式,将多模态图像融合中的特征分解重构为积分驱动的一维信号优化问题,通过两阶段自监督学习实现稳定优化,并在代表性任务上取得最先进性能。

AI 中文摘要

多模态图像融合(MMIF)旨在将来自不同模态的互补信息整合成高质量的融合图像,并支持下游任务。近年来,特征分解通过将源图像分离为共有特征和模态特有的独特特征,已成为一种重要范式。然而,现有方法缺乏明确的监督,因为无法获得真值(GT)分解特征图。它们通常将多个图像级指标组合作为损失,这些指标本质上不完整且可能相互冲突,因为每个像素耦合了纹理、边缘和轮廓等属性。为解决此问题,我们提出了一种一维信号级自监督特征分解范式。我们的核心见解是将特征分解从模糊的二维图像级监督重新表述为积分驱动的一维信号级优化问题。这种目标级重构利用一维信号形式来计算积分约束。分解器通过共有信号与原始信号之间的积分面积进行优化,从而在明确的优化目标下实现更稳定的优化。我们的模型遵循两阶段自监督学习(SSL)框架。第一阶段设计了双预文本任务:信号级积分驱动分解和图像级结构保持重建。第二阶段融合独特特征,并将其与共有特征结合以重建融合图像。在代表性MMIF任务上的实验表明,我们的方法达到了最先进的(SOTA)性能。代码:此HTTP URL。

英文摘要

Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a high-quality fused image and support downstream tasks. Recently, feature decomposition has become an important paradigm by separating source images into common and modality-specific unique features. However, existing methods lack clear supervision because ground-truth (GT) decomposition feature maps are unavailable. They usually combine multiple image-level metrics as losses, which are inherently incomplete and may conflict since each pixel couples attributes such as texture, edge, and contour. To address this, we propose a 1D signal-level self-supervised feature decomposition paradigm. Our core insight is to reformulate feature decomposition from unclear 2D image-level supervision into an integral-driven 1D signal-level optimization problem. This objective-level reformulation uses the 1D signal form to compute the integral constraint. The decomposer is optimized by the integral area between common and original signals, enabling more stable optimization with a clear optimization objective. Our model follows a two-stage SSL framework. Stage I designs dual pretext tasks for integral-driven decomposition at the signal level and structure-preserving reconstruction at the image level. Stage II fuses unique features and combines them with common features to reconstruct the fused image. Experiments on representative MMIF tasks show state-of-the-art (SOTA) performance. Code: github.com/Wangjiayu0512/SIDFusion.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑