arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

至 收录 4712 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM
2508.15505 2025-08-22 cs.CV 92%

Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion

Mengyu Wang, Zhenyu Liu, Kun Li, Yu Wang, Yuwei Wang, Yanyan Wei, Fei Wang

机构 * Key Laboratory of Opto-Electronic Information Science and Technology of Jiangxi Province, Nanchang Hangkong University(江西省光电信息科学与技术重点实验室,南昌航空大学) ReLER, CCAI, Zhejiang University(ReLER、CCAI、浙江大学) College of Engineering, Anhui Agricultural University(安徽农业大学工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 通用Image Fusion :image fusion(title,abstract);multimodal image fusion(title,abstract);multimodal fusion(abstract);multi-focus(abstract)

Comments Accepted by IEEE Transactions on Multimedia

详情
英文摘要

Multimodal Image Fusion (MMIF) aims to integrate complementary information from different imaging modalities to overcome the limitations of individual sensors. It enhances image quality and facilitates downstream applications such as remote sensing, medical diagnostics, and robotics. Despite significant advancements, current MMIF methods still face challenges such as modality misalignment, high-frequency detail destruction, and task-specific limitations. To address these challenges, we propose AdaSFFuse, a novel framework for task-generalized MMIF through adaptive cross-domain co-fusion learning. AdaSFFuse introduces two key innovations: the Adaptive Approximate Wavelet Transform (AdaWAT) for frequency decoupling, and the Spatial-Frequency Mamba Blocks for efficient multimodal fusion. AdaWAT adaptively separates the high- and low-frequency components of multimodal images from different scenes, enabling fine-grained extraction and alignment of distinct frequency characteristics for each modality. The Spatial-Frequency Mamba Blocks facilitate cross-domain fusion in both spatial and frequency domains, enhancing this process. These blocks dynamically adjust through learnable mappings to ensure robust fusion across diverse modalities. By combining these components, AdaSFFuse improves the alignment and integration of multimodal features, reduces frequency loss, and preserves critical details. Extensive experiments on four MMIF tasks -- Infrared-Visible Image Fusion (IVF), Multi-Focus Image Fusion (MFF), Multi-Exposure Image Fusion (MEF), and Medical Image Fusion (MIF) -- demonstrate AdaSFFuse's superior fusion performance, ensuring both low computational cost and a compact network, offering a strong balance between performance and efficiency. The code will be publicly available at https://github.com/Zhen-yu-Liu/AdaSFFuse.

URL PDF HTML 收藏
2311.01886 2024-02-01 cs.CV 91%

Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image Fusion

Xilai Li, Xiaosong Li, Tao Ye, Xiaoqi Cheng, Wuyang Liu, Haishu Tan

专题命中 通用Image Fusion :image fusion(title,abstract);multi-modal image fusion(title,abstract);multi-focus(title);分类 cs.CV

Comments Accepted to IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2024

详情
英文摘要

Multi-modal image fusion (MMIF) integrates valuable information from different modality images into a fused one. However, the fusion of multiple visible images with different focal regions and infrared images is a unprecedented challenge in real MMIF applications. This is because of the limited depth of the focus of visible optical lenses, which impedes the simultaneous capture of the focal information within the same scene. To address this issue, in this paper, we propose a MMIF framework for joint focused integration and modalities information extraction. Specifically, a semi-sparsity-based smoothing filter is introduced to decompose the images into structure and texture components. Subsequently, a novel multi-scale operator is proposed to fuse the texture components, capable of detecting significant information by considering the pixel focus attributes and relevant data from various modal images. Additionally, to achieve an effective capture of scene luminance and reasonable contrast maintenance, we consider the distribution of energy information in the structural components in terms of multi-directional frequency variance and information entropy. Extensive experiments on existing MMIF datasets, as well as the object detection and depth estimation tasks, consistently demonstrate that the proposed algorithm can surpass the state-of-the-art methods in visual perception and quantitative evaluation. The code is available at https://github.com/ixilai/MFIF-MMIF.

URL PDF HTML 收藏
2509.09456 2025-09-12 cs.CV 90%

FlexiD-Fuse: Flexible number of inputs multi-modal medical image fusion based on diffusion model

Yushen Xu, Xiaosong Li, Yuchun Wang, Xiaoqi Cheng, Huafeng Li, Haishu Tan

机构 * School of Physics(物理学院) Optoelectronic Engineering, Foshan University(光电工程学院,佛山大学) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology(粤港澳联合智能微纳光电技术实验室) Guangdong Provincial Key Laboratory of Industrial Intelligent Inspection Technology(广东省工业智能检测技术重点实验室) School of Information Engineering(信息工程学院) Automation, Kunming University of Science(自动化系,昆明理工大学)

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);multi-focus(abstract);multi-exposure(abstract)

Journal ref Expert Systems with Applications, 2025: 128895

详情
英文摘要

Different modalities of medical images provide unique physiological and anatomical information for diseases. Multi-modal medical image fusion integrates useful information from different complementary medical images with different modalities, producing a fused image that comprehensively and objectively reflects lesion characteristics to assist doctors in clinical diagnosis. However, existing fusion methods can only handle a fixed number of modality inputs, such as accepting only two-modal or tri-modal inputs, and cannot directly process varying input quantities, which hinders their application in clinical settings. To tackle this issue, we introduce FlexiD-Fuse, a diffusion-based image fusion network designed to accommodate flexible quantities of input modalities. It can end-to-end process two-modal and tri-modal medical image fusion under the same weight. FlexiD-Fuse transforms the diffusion fusion problem, which supports only fixed-condition inputs, into a maximum likelihood estimation problem based on the diffusion process and hierarchical Bayesian modeling. By incorporating the Expectation-Maximization algorithm into the diffusion sampling iteration process, FlexiD-Fuse can generate high-quality fused images with cross-modal information from source images, independently of the number of input images. We compared the latest two and tri-modal medical image fusion methods, tested them on Harvard datasets, and evaluated them using nine popular metrics. The experimental results show that our method achieves the best performance in medical image fusion with varying inputs. Meanwhile, we conducted extensive extension experiments on infrared-visible, multi-exposure, and multi-focus image fusion tasks with arbitrary numbers, and compared them with the perspective SOTA methods. The results of the extension experiments consistently demonstrate the effectiveness and superiority of our method.

URL PDF HTML 收藏
2007.13538 2020-07-28 cs.CV eess.IV 90%

A Novel adaptive optimization of Dual-Tree Complex Wavelet Transform for Medical Image Fusion

T. Deepika, G. Karpaga Kannan

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);multimodal image fusion(abstract);分类 cs.CV、eess.IV

Comments Conference on Computing Communication and Signal Processing. arXiv admin note: text overlap with arXiv:2007.11488

详情
英文摘要

In recent years, many research achievements are made in the medical image fusion field. Fusion is basically extraction of best of inputs and conveying it to the output. Medical Image fusion means that several of various modality image information is comprehended together to form one image to express its information. The aim of image fusion is to integrate complementary and redundant information. In this paper, a multimodal image fusion algorithm based on the dual-tree complex wavelet transform (DT-CWT) and adaptive particle swarm optimization (APSO) is proposed. Fusion is achieved through the formation of a fused pyramid using the DTCWT coefficients from the decomposed pyramids of the source images. The coefficients are fused by the weighted average method based on pixels, and the weights are estimated by the APSO to gain optimal fused images. The fused image is obtained through conventional inverse dual-tree complex wavelet transform reconstruction process. Experiment results show that the proposed method based on adaptive particle swarm optimization algorithm is remarkably better than the method based on particle swarm optimization. The resulting fused images are compared visually and through benchmarks such as Entropy (E), Peak Signal to Noise Ratio, (PSNR), Root Mean Square Error (RMSE), Standard deviation (SD) and Structure Similarity Index Metric (SSIM) computations.

URL PDF HTML 收藏
2603.24296 2026-03-26 cs.CV 89%

AMIF: Authorizable Medical Image Fusion Model with Built-in Authentication

AMIF:具有内置认证的可授权医学图像融合模型

Jie Song, Jun Jia, Wei Sun, Wangqiu Zhou, Tao Tan, Guangtao Zhai

机构 * Macao Polytechnic University(澳门理工学院) Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学) Hefei University of Technology(合肥工业大学)

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);multimodal image fusion(abstract);分类 cs.CV

AI总结 本文提出AMIF,首个具有内置认证的医学图像融合模型,通过授权访问控制保护知识产权,防止推理泄露。

详情
AI中文摘要

多模态图像融合能够实现精确的病变定位和特征描述,从而提高诊断准确性,增强临床决策并推动其在医学影像研究中的重要性。强大的多模态图像融合模型依赖于高质量、临床代表性的多模态训练数据和精心设计的模型架构。因此,此类专业放射组学模型的开发是标准化采集、临床专业知识和算法设计能力的协作成果,需要保护相关知识产权。然而,当前多模态图像融合模型生成融合输出,但缺乏内置机制来保护知识产权,导致通过推理泄露暴露了专有模型知识和敏感训练数据。例如,恶意用户可以利用融合输出和模型蒸馏或其他基于推理的逆向工程技术来近似专有模型的融合性能。为了解决这个问题,我们提出了AMIF,首个具有内置认证的医学图像融合模型,将授权访问控制整合到图像融合目标中。对于未经授权的使用,AMIF将显式可见的版权标识嵌入融合结果中。相反,通过成功的关键基于认证可以访问高质量的融合结果。

英文摘要

Multimodal image fusion enables precise lesion localization and characterization for accurate diagnosis, thereby strengthening clinical decision-making and driving its growing prominence in medical imaging research. A powerful multimodal image fusion model relies on high-quality, clinically representative multimodal training data and a rigorously engineered model architecture. Therefore, the development of such professional radiomics models represents a collaborative achievement grounded in standardized acquisition, clinical-specific expertise, and algorithmic design proficiency, which necessitates protection of associated intellectual property rights. However, current multimodal image fusion models generate fused outputs without built-in mechanisms to safeguard intellectual property rights, inadvertently exposing proprietary model knowledge and sensitive training data through inference leakage. For example, malicious users can exploit fusion outputs and model distillation or other inference-based reverse engineering techniques to approximate the fusion performance of proprietary models. To address this issue, we propose AMIF, the first Authorizable Medical Image Fusion model with built-in authentication, which integrates authorization access control into the image fusion objective. For unauthorized usage, AMIF embeds explicit and visible copyright identifiers into fusion results. In contrast, high-quality fusion results are accessible upon successful key-based authentication.

URL PDF HTML 收藏
2411.10036 2025-12-12 cs.CV cs.AI 89%

Rethinking Normalization Strategies and Convolutional Kernels for Multimodal Image Fusion

重新思考归一化策略和卷积核在多模态图像融合中的作用

Dan He, Guofen Wang, Weisheng Li, Yucheng Shu, Wenbo Li, Lijian Yang, Yuping Huang, Feiyan Li

机构 * Chongqing University of Posts and Telecommunications(重庆邮电大学) Chongqing Normal University(重庆师范大学)

专题命中 通用Image Fusion :image fusion(title,abstract);multimodal image fusion(title,abstract);information fusion(abstract);分类 cs.CV

AI总结 本文提出了一种改进的UNet架构,通过混合实例和组归一化以及多路径自适应融合模块,提升多模态图像融合的性能和细节保留能力。

详情
AI中文摘要

多模态图像融合(MMIF)整合不同模态的信息以获得综合图像,有助于后续任务。然而,现有研究侧重于互补信息融合和训练策略,忽略了归一化和卷积核等基础组件的关键作用。我们重新评估了UNet架构用于端到端的MMIF,发现广泛使用的批量归一化通过平滑关键稀疏特征限制了性能。为了解决这个问题,我们提出了一种实例和组归一化的混合方法,以保持样本独立性和强化内在特征相关性。关键的是,这种策略促进了更丰富的特征图,使大核卷积能够充分利用其感受野,增强细节保留。此外,所提出的多路径自适应融合模块动态校准不同尺度和感受野的特征,确保有效信息传输。我们的方法在MSRS、M³FD、TNO和哈佛数据集上实现了SOTA目标性能,生成视觉上更清晰的显著对象和病变区域。值得注意的是,它在红外图像上的MSRS分割mIoU提高了8.1%。这种性能源于归一化和卷积核的协同设计,保留了关键稀疏特征。代码可在https://github.com/HeDan-11/LKC-FUNet上获得。

英文摘要

Multimodal image fusion (MMIF) integrates information from different modalities to obtain a comprehensive image, aiding downstream tasks. However, existing research focuses on complementary information fusion and training strategies, overlooking the critical role of underlying architectural components like normalization and convolution kernels. We reevaluate the UNet architecture for end-to-end MMIF, identifying that widely used batch normalization limits performance by smoothing crucial sparse features. To address this, we propose a hybrid of instance and group normalization to maintain sample independence and reinforce intrinsic feature correlations. Crucially, this strategy facilitates richer feature maps, enabling large kernel convolution to fully leverage its receptive field, enhancing detail preservation. Furthermore, the proposed multi-path adaptive fusion module dynamically calibrates features from varying scales and receptive fields, ensuring effective information transfer. Our method achieves SOTA objective performance on MSRS, M$^3$FD, TNO, and Harvard datasets, producing visually clearer salient objects and lesion areas. Notably, it improves MSRS segmentation mIoU by 8.1\% over the infrared image. This performance stems from a synergistic design of normalization and convolution kernels, which preserves critical sparse features. The code is available at https://github.com/HeDan-11/LKC-FUNet.

URL PDF HTML 收藏
2410.23905 2024-11-01 cs.CV 89%

Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model

Hao Zhang, Lei Cao, Jiayi Ma

专题命中 通用Image Fusion :image fusion(title,abstract);multi-modal image fusion(title,abstract);information fusion(abstract);分类 cs.CV

Comments Accepted by the 38th Conference on Neural Information Processing Systems (NeurIPS 2024)

详情
英文摘要

Existing multi-modal image fusion methods fail to address the compound degradations presented in source images, resulting in fusion images plagued by noise, color bias, improper exposure, \textit{etc}. Additionally, these methods often overlook the specificity of foreground objects, weakening the salience of the objects of interest within the fused images. To address these challenges, this study proposes a novel interactive multi-modal image fusion framework based on the text-modulated diffusion model, called Text-DiFuse. First, this framework integrates feature-level information integration into the diffusion process, allowing adaptive degradation removal and multi-modal information fusion. This is the first attempt to deeply and explicitly embed information fusion within the diffusion process, effectively addressing compound degradation in image fusion. Second, by embedding the combination of the text and zero-shot location model into the diffusion fusion process, a text-controlled fusion re-modulation strategy is developed. This enables user-customized text control to improve fusion performance and highlight foreground objects in the fused images. Extensive experiments on diverse public datasets show that our Text-DiFuse achieves state-of-the-art fusion performance across various scenarios with complex degradation. Moreover, the semantic segmentation experiment validates the significant enhancement in semantic performance achieved by our text-controlled fusion re-modulation strategy. The code is publicly available at https://github.com/Leiii-Cao/Text-DiFuse.

URL PDF HTML 收藏
2310.05462 2023-10-25 cs.CV 89%

AdaFuse: Adaptive Medical Image Fusion Based on Spatial-Frequential Cross Attention

Xianming Gu, Lihui Wang, Zeyu Deng, Ying Cao, Xingyu Huang, Yue-min Zhu

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);information fusion(abstract);分类 cs.CV

详情
英文摘要

Multi-modal medical image fusion is essential for the precise clinical diagnosis and surgical navigation since it can merge the complementary information in multi-modalities into a single image. The quality of the fused image depends on the extracted single modality features as well as the fusion rules for multi-modal information. Existing deep learning-based fusion methods can fully exploit the semantic features of each modality, they cannot distinguish the effective low and high frequency information of each modality and fuse them adaptively. To address this issue, we propose AdaFuse, in which multimodal image information is fused adaptively through frequency-guided attention mechanism based on Fourier transform. Specifically, we propose the cross-attention fusion (CAF) block, which adaptively fuses features of two modalities in the spatial and frequency domains by exchanging key and query values, and then calculates the cross-attention scores between the spatial and frequency features to further guide the spatial-frequential information fusion. The CAF block enhances the high-frequency features of the different modalities so that the details in the fused images can be retained. Moreover, we design a novel loss function composed of structure loss and content loss to preserve both low and high frequency information. Extensive comparison experiments on several datasets demonstrate that the proposed method outperforms state-of-the-art methods in terms of both visual quality and quantitative metrics. The ablation experiments also validate the effectiveness of the proposed loss and fusion strategy.

URL PDF HTML 收藏
2603.23272 2026-03-25 cs.CV cs.MM 88%

Multi-Modal Image Fusion via Intervention-Stable Feature Learning

多模态图像融合 via 干预稳定的特征学习

Xue Wang, Zheng Guan, Wenhua Qian, Chengchao Wang, Runzhuo Ma

机构 * School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) School of Artificial Intelligence, Nanyang Normal University(南阳师范学院人工智能学院) Department of Electrical and Electronic Engineering, Hong Kong Polytechnic University(香港理工大学电子与电气工程系)

专题命中 通用Image Fusion :image fusion(title,abstract);multi-modal image fusion(title,abstract);分类 cs.CV、cs.MM

AI总结 本文提出基于因果原则的干预框架,通过三种干预策略探索模态关系,设计因果特征整合器捕捉稳健的模态依赖而非虚假关联。

Comments Accpted by CVPR 2026

详情
AI中文摘要

多模态图像融合将不同模态的互补信息整合为统一表示。当前方法主要优化模态间的统计相关性,常捕捉数据集诱导的虚假关联,在分布偏移下劣化。本文提出基于因果原则的干预框架,通过互补遮蔽、随机遮蔽和模态dropout三种策略探索模态关系,设计因果特征整合器学习干预稳定的特征,从而捕捉稳健的模态依赖而非虚假关联。实验表明,本文方法在公开基准和下游高阶视觉任务中均取得SOTA性能。

英文摘要

Multi-modal image fusion integrates complementary information from different modalities into a unified representation. Current methods predominantly optimize statistical correlations between modalities, often capturing dataset-induced spurious associations that degrade under distribution shifts. In this paper, we propose an intervention-based framework inspired by causal principles to identify robust cross-modal dependencies. Drawing insights from Pearl's causal hierarchy, we design three principled intervention strategies to probe different aspects of modal relationships: i) complementary masking with spatially disjoint perturbations tests whether modalities can genuinely compensate for each other's missing information, ii) random masking of identical regions identifies feature subsets that remain informative under partial observability, and iii) modality dropout evaluates the irreplaceable contribution of each modality. Based on these interventions, we introduce a Causal Feature Integrator (CFI) that learns to identify and prioritize intervention-stable features maintaining importance across different perturbation patterns through adaptive invariance gating, thereby capturing robust modal dependencies rather than spurious correlations. Extensive experiments demonstrate that our method achieves SOTA performance on both public benchmarks and downstream high-level vision tasks.

URL PDF HTML 收藏
2602.04405 2026-02-05 cs.CV cs.MM 88%

Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion

交互式空间-频率融合Mamba用于多模态图像融合

Yixin Zhu, Long Lv, Pingping Zhang, Xuehu Liu, Tongdan Tang, Feng Tian, Weibing Sun, Huchuan Lu

机构 * School of Future Technology, Dalian University of Technology and the Key Laboratory of Data Science and Smart Education (Hainan Normal University), Ministry of Education(未来技术学院,大连理工大学和数据科学与智能教育关键实验室(海南师范大学),教育部) Affiliated Zhongshan Hospital of Dalian University(大连大学附属中山医院) School of Computer Science and Artificial Intelligence, Wuhan University of Technology(计算机科学与人工智能学院,武汉理工大学) Central Hospital of Dalian University of Technology(大连理工大学中心医院) School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学)

专题命中 通用Image Fusion :image fusion(title,abstract);multi-modal image fusion(title,abstract);分类 cs.CV、cs.MM

AI总结 本文提出交互式空间-频率融合Mamba框架,通过多尺度频率融合和交互式融合提升多模态图像融合性能。

Comments This work is accepted by IEEE Transactions on Image Processing. More modifications may be performed

详情
AI中文摘要

多模态图像融合(MMIF)旨在将不同模态的图像结合起来,生成融合图像,保留纹理细节并保持重要信息。最近,一些MMIF方法结合了频率域信息以增强空间特征。然而,这些方法通常依赖于简单的串行或并行空间-频率融合,没有交互。在本文中,我们提出了一种新的交互式空间-频率融合Mamba(ISFM)框架用于MMIF。具体来说,我们首先使用模态特定提取器(MSE)从不同模态中提取特征。它在图像上建模长距离依赖关系,具有线性计算复杂度。为了有效地利用频率信息,我们随后提出多尺度频率融合(MFF)。它在多个尺度上自适应地整合低频和高频成分,使频率特征具有鲁棒的表示。更重要的是,我们进一步提出交互式空间-频率融合(ISF)。它将频率特征融入到跨模态的空间特征中,增强互补的表示。在六个MMIF数据集上进行了广泛的实验。实验结果表明,我们的ISFM比其他最先进的方法在性能上更优。源代码可在https://github.com/Namn23/ISFM上获得。

英文摘要

Multi-Modal Image Fusion (MMIF) aims to combine images from different modalities to produce fused images, retaining texture details and preserving significant information. Recently, some MMIF methods incorporate frequency domain information to enhance spatial features. However, these methods typically rely on simple serial or parallel spatial-frequency fusion without interaction. In this paper, we propose a novel Interactive Spatial-Frequency Fusion Mamba (ISFM) framework for MMIF. Specifically, we begin with a Modality-Specific Extractor (MSE) to extract features from different modalities. It models long-range dependencies across the image with linear computational complexity. To effectively leverage frequency information, we then propose a Multi-scale Frequency Fusion (MFF). It adaptively integrates low-frequency and high-frequency components across multiple scales, enabling robust representations of frequency features. More importantly, we further propose an Interactive Spatial-Frequency Fusion (ISF). It incorporates frequency features to guide spatial features across modalities, enhancing complementary representations. Extensive experiments are conducted on six MMIF datasets. The experimental results demonstrate that our ISFM can achieve better performances than other state-of-the-art methods. The source code is available at https://github.com/Namn23/ISFM.

URL PDF HTML 收藏
2412.08050 2024-12-16 eess.IV cs.CV cs.LG 88%

BSAFusion: A Bidirectional Stepwise Feature Alignment Network for Unaligned Medical Image Fusion

Huafeng Li, Dayong Su, Qing Cai, Yafei Zhang

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);分类 cs.CV、eess.IV

Comments Accepted by AAAI2025

详情
英文摘要

If unaligned multimodal medical images can be simultaneously aligned and fused using a single-stage approach within a unified processing framework, it will not only achieve mutual promotion of dual tasks but also help reduce the complexity of the model. However, the design of this model faces the challenge of incompatible requirements for feature fusion and alignment; specifically, feature alignment requires consistency among corresponding features, whereas feature fusion requires the features to be complementary to each other. To address this challenge, this paper proposes an unaligned medical image fusion method called Bidirectional Stepwise Feature Alignment and Fusion (BSFA-F) strategy. To reduce the negative impact of modality differences on cross-modal feature matching, we incorporate the Modal Discrepancy-Free Feature Representation (MDF-FR) method into BSFA-F. MDF-FR utilizes a Modality Feature Representation Head (MFRH) to integrate the global information of the input image. By injecting the information contained in MFRH of the current image into other modality images, it effectively reduces the impact of modality differences on feature alignment while preserving the complementary information carried by different images. In terms of feature alignment, BSFA-F employs a bidirectional stepwise alignment deformation field prediction strategy based on the path independence of vector displacement between two points. This strategy solves the problem of large spans and inaccurate deformation field prediction in single-step alignment. Finally, Multi-Modal Feature Fusion block achieves the fusion of aligned features. The experimental results across multiple datasets demonstrate the effectiveness of our method. The source code is available at https://github.com/slrl123/BSAFusion.

URL PDF HTML 收藏
2411.11799 2024-11-19 eess.IV cs.AI cs.CV 88%

Edge-Enhanced Dilated Residual Attention Network for Multimodal Medical Image Fusion

Meng Zhou, Yuxuan Zhang, Xiaolan Xu, Jiayi Wang, Farzad Khalvati

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);分类 cs.CV、eess.IV

Comments An extended version of the paper accepted at IEEE BIBM 2024

详情
英文摘要

Multimodal medical image fusion is a crucial task that combines complementary information from different imaging modalities into a unified representation, thereby enhancing diagnostic accuracy and treatment planning. While deep learning methods, particularly Convolutional Neural Networks (CNNs) and Transformers, have significantly advanced fusion performance, some of the existing CNN-based methods fall short in capturing fine-grained multiscale and edge features, leading to suboptimal feature integration. Transformer-based models, on the other hand, are computationally intensive in both the training and fusion stages, making them impractical for real-time clinical use. Moreover, the clinical application of fused images remains unexplored. In this paper, we propose a novel CNN-based architecture that addresses these limitations by introducing a Dilated Residual Attention Network Module for effective multiscale feature extraction, coupled with a gradient operator to enhance edge detail learning. To ensure fast and efficient fusion, we present a parameter-free fusion strategy based on the weighted nuclear norm of softmax, which requires no additional computations during training or inference. Extensive experiments, including a downstream brain tumor classification task, demonstrate that our approach outperforms various baseline methods in terms of visual quality, texture preservation, and fusion speed, making it a possible practical solution for real-world clinical applications. The code will be released at https://github.com/simonZhou86/en_dran.

URL PDF HTML 收藏
2404.17357 2024-10-16 eess.IV cs.CV 88%

Simultaneous Tri-Modal Medical Image Fusion and Super-Resolution using Conditional Diffusion Model

Yushen Xu, Xiaosong Li, Yuchan Jie, Haishu Tan

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);分类 cs.CV、eess.IV

Comments Accepted by MICCAI 2024

Journal ref International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer Nature Switzerland, 2024: 635-645

详情
英文摘要

In clinical practice, tri-modal medical image fusion, compared to the existing dual-modal technique, can provide a more comprehensive view of the lesions, aiding physicians in evaluating the disease's shape, location, and biological activity. However, due to the limitations of imaging equipment and considerations for patient safety, the quality of medical images is usually limited, leading to sub-optimal fusion performance, and affecting the depth of image analysis by the physician. Thus, there is an urgent need for a technology that can both enhance image resolution and integrate multi-modal information. Although current image processing methods can effectively address image fusion and super-resolution individually, solving both problems synchronously remains extremely challenging. In this paper, we propose TFS-Diff, a simultaneously realize tri-modal medical image fusion and super-resolution model. Specially, TFS-Diff is based on the diffusion model generation of a random iterative denoising process. We also develop a simple objective function and the proposed fusion super-resolution loss, effectively evaluates the uncertainty in the fusion and ensures the stability of the optimization process. And the channel attention module is proposed to effectively integrate key information from different modalities for clinical diagnosis, avoiding information loss caused by multiple image processing. Extensive experiments on public Harvard datasets show that TFS-Diff significantly surpass the existing state-of-the-art methods in both quantitative and visual evaluations. Code is available at https://github.com/XylonXu01/TFS-Diff.

URL PDF HTML 收藏
2206.15179 2024-07-09 eess.IV cs.CV cs.LG 88%

D2-LRR: A Dual-Decomposed MDLatLRR Approach for Medical Image Fusion

Xu Song, Tianyu Shen, Hui Li, Xiao-Jun Wu

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);分类 cs.CV、eess.IV

Comments There are some errors that need to be corrected

详情
英文摘要

In image fusion tasks, an ideal image decomposition method can bring better performance. MDLatLRR has done a great job in this aspect, but there is still exist some space for improvement. Considering that MDLatLRR focuses solely on the detailed parts (salient features) extracted from input images via latent low-rank representation (LatLRR), the basic parts (principal features) extracted by LatLRR are not fully utilized. Therefore, we introduced an enhanced multi-level decomposition method named dual-decomposed MDLatLRR (D2-LRR) which effectively analyzes and utilizes all image features extracted through LatLRR. Specifically, color images are converted into YUV color space and grayscale images, and the Y-channel and grayscale images are input into the trained parameters of LatLRR to obtain the detailed parts containing four rounds of decomposition and the basic parts. Subsequently, the basic parts are fused using an average strategy, while the detail part is fused using kernel norm operation. The fused image is ultimately transformed back into an RGB image, resulting in the final fusion output. We apply D2-LRR to medical image fusion tasks. The detailed parts are fused employing a nuclear-norm operation, while the basic parts are fused using an average strategy. Comparative analyses among existing methods showcase that our proposed approach attains cutting-edge fusion performance in both objective and subjective assessments.

URL PDF HTML 收藏
2310.11896 2023-10-19 eess.IV cs.CV cs.LG 88%

A New Multimodal Medical Image Fusion based on Laplacian Autoencoder with Channel Attention

Payal Wankhede, Manisha Das, Deep Gupta, Petia Radeva, Ashwini M Bakde

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);分类 cs.CV、eess.IV

Comments 10 pages, 6 figures, % tables

详情
英文摘要

Medical image fusion combines the complementary information of multimodal medical images to assist medical professionals in the clinical diagnosis of patients' disorders and provide guidance during preoperative and intra-operative procedures. Deep learning (DL) models have achieved end-to-end image fusion with highly robust and accurate fusion performance. However, most DL-based fusion models perform down-sampling on the input images to minimize the number of learnable parameters and computations. During this process, salient features of the source images become irretrievable leading to the loss of crucial diagnostic edge details and contrast of various brain tissues. In this paper, we propose a new multimodal medical image fusion model is proposed that is based on integrated Laplacian-Gaussian concatenation with attention pooling (LGCA). We prove that our model preserves effectively complementary information and important tissue structures.

URL PDF HTML 收藏
1912.07959 2022-02-01 cs.CV eess.IV 88%

Multi-focus Image Fusion Based on Similarity Characteristics

Ya-Qiong Zhang, Xiao-Jun Wu, Hui Li

专题命中 通用Image Fusion :image fusion(title,abstract);multi-focus(title,abstract);分类 cs.CV、eess.IV

Comments 7 pages, 10 figures, 4 tables

详情
英文摘要

A novel multi-focus image fusion algorithm performed in spatial domain based on similarity characteristics is proposed incorporating with region segmentation. In this paper, a new similarity measure is developed based on the structural similarity (SSIM) index, which is more suitable for multi-focus image segmentation. Firstly, the SSNSIM map is calculated between two input images. Then we segment the SSNSIM map using watershed method, and merge the small homogeneous regions with fuzzy c-means clustering algorithm (FCM). For three source images, a joint region segmentation method based on segmentation of two images is used to obtain the final segmentation result. Finally, the corresponding segmented regions of the source images are fused according to their average gradient. The performance of the image fusion method is evaluated by several criteria including spatial frequency, average gradient, entropy, edge retention etc. The evaluation results indicate that the proposed method is effective and has good visual perception.

URL PDF HTML 收藏
2012.14678 2021-02-02 cs.CV eess.IV 88%

Towards Reducing Severe Defocus Spread Effects for Multi-Focus Image Fusion via an Optimization Based Strategy

Shuang Xu, Lizhen Ji, Zhe Wang, Pengfei Li, Kai Sun, Chunxia Zhang, Jiangshe Zhang

专题命中 通用Image Fusion :image fusion(title,abstract);multi-focus(title,abstract);分类 cs.CV、eess.IV

Journal ref IEEE Transactions on Computational Imaging, vol. 6, pp. 1561-1570, 2020

详情
英文摘要

Multi-focus image fusion (MFF) is a popular technique to generate an all-in-focus image, where all objects in the scene are sharp. However, existing methods pay little attention to defocus spread effects of the real-world multi-focus images. Consequently, most of the methods perform badly in the areas near focus map boundaries. According to the idea that each local region in the fused image should be similar to the sharpest one among source images, this paper presents an optimization-based approach to reduce defocus spread effects. Firstly, a new MFF assessmentmetric is presented by combining the principle of structure similarity and detected focus maps. Then, MFF problem is cast into maximizing this metric. The optimization is solved by gradient ascent. Experiments conducted on the real-world dataset verify superiority of the proposed model. The codes are available at https://github.com/xsxjtu/MFF-SSIM.

URL PDF HTML 收藏
2009.13615 2020-10-06 eess.IV cs.CV 88%

Multi-focus Image Fusion for Visual Sensor Networks

Milad Abdollahzadeh, Touba Malekzadeh, Hadi Seyedarabi

专题命中 通用Image Fusion :image fusion(title,abstract);multi-focus(title,abstract);分类 cs.CV、eess.IV

Comments 5 pages

详情
英文摘要

Image fusion in visual sensor networks (VSNs) aims to combine information from multiple images of the same scene in order to transform a single image with more information. Image fusion methods based on discrete cosine transform (DCT) are less complex and time-saving in DCT based standards of image and video which makes them more suitable for VSN applications. In this paper, an efficient algorithm for the fusion of multi-focus images in the DCT domain is proposed. The Sum of modified laplacian (SML) of corresponding blocks of source images is used as a contrast criterion and blocks with the larger value of SML are absorbed to output images. The experimental results on several images show the improvement of the proposed algorithm in terms of both subjective and objective quality of fused image relative to other DCT based techniques.

URL PDF HTML 收藏
2007.15156 2020-07-31 cs.CV eess.IV 88%

Benchmarking and Comparing Multi-exposure Image Fusion Algorithms

Xingchen Zhang

专题命中 通用Image Fusion :image fusion(title,abstract);multi-exposure(title,abstract);分类 cs.CV、eess.IV

Comments 24 pages, 5 figures, 4 tables

详情
英文摘要

Multi-exposure image fusion (MEF) is an important area in computer vision and has attracted increasing interests in recent years. Apart from conventional algorithms, deep learning techniques have also been applied to multi-exposure image fusion. However, although much efforts have been made on developing MEF algorithms, the lack of benchmark makes it difficult to perform fair and comprehensive performance comparison among MEF algorithms, thus significantly hindering the development of this field. In this paper, we fill this gap by proposing a benchmark for multi-exposure image fusion (MEFB) which consists of a test set of 100 image pairs, a code library of 16 algorithms, 20 evaluation metrics, 1600 fused images and a software toolkit. To the best of our knowledge, this is the first benchmark in the field of multi-exposure image fusion. Extensive experiments have been conducted using MEFB for comprehensive performance evaluation and for identifying effective algorithms. We expect that MEFB will serve as an effective platform for researchers to compare performances and investigate MEF algorithms.

URL PDF HTML 收藏
2002.04780 2020-02-13 cs.CV cs.MM 88%

MFFW: A new dataset for multi-focus image fusion

Shuang Xu, Xiaoli Wei, Chunxia Zhang, Junmin Liu, Jiangshe Zhang

专题命中 通用Image Fusion :image fusion(title,abstract);multi-focus(title,abstract);分类 cs.CV、cs.MM

详情
英文摘要

Multi-focus image fusion (MFF) is a fundamental task in the field of computational photography. Current methods have achieved significant performance improvement. It is found that current methods are evaluated on simulated image sets or Lytro dataset. Recently, a growing number of researchers pay attention to defocus spread effect, a phenomenon of real-world multi-focus images. Nonetheless, defocus spread effect is not obvious in simulated or Lytro datasets, where popular methods perform very similar. To compare their performance on images with defocus spread effect, this paper constructs a new dataset called MFF in the wild (MFFW). It contains 19 pairs of multi-focus images collected on the Internet. We register all pairs of source images, and provide focus maps and reference images for part of pairs. Compared with Lytro dataset, images in MFFW significantly suffer from defocus spread effect. In addition, the scenes of MFFW are more complex. The experiments demonstrate that most state-of-the-art methods on MFFW dataset cannot robustly generate satisfactory fusion images. MFFW can be a new baseline dataset to test whether an MMF algorithm is able to deal with defocus spread effect.

URL PDF HTML 收藏
1906.00225 2019-12-12 eess.IV cs.CV 88%

A Semantic-based Medical Image Fusion Approach

Fanda Fan, Yunyou Huang, Lei Wang, Xingwang Xiong, Zihan Jiang, Zhifei Zhang, Jianfeng Zhan

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);分类 cs.CV、eess.IV

详情
英文摘要

It is necessary for clinicians to comprehensively analyze patient information from different sources. Medical image fusion is a promising approach to providing overall information from medical images of different modalities. However, existing medical image fusion approaches ignore the semantics of images, making the fused image difficult to understand. In this work, we propose a new evaluation index to measure the semantic loss of fused image, and put forward a Fusion W-Net (FW-Net) for multimodal medical image fusion. The experimental results are promising: the fused image generated by our approach greatly reduces the semantic information loss, and has better visual effects in contrast to five state-of-art approaches. Our approach and tool have great potential to be applied in the clinical setting.

URL PDF HTML 收藏
2509.09427 2025-09-12 cs.CV 88%

FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

Yuchan Jie, Yushen Xu, Xiaosong Li, Fuqiang Zhou, Jianming Lv, Huafeng Li

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Physics and Optoelectronic Engineering, Foshan University(佛山大学物理与光电工程学院) School of Instrumentation Science and Optelectronics Engineering, Beihang University(北航仪器科学与光电工程学院) School of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院)

专题命中 通用Image Fusion :image fusion(title,abstract);multimodal image fusion(title);information fusion(abstract,journal_ref);分类 cs.CV

Journal ref Information Fusion, 2025, 121: 103146

详情
英文摘要

As an influential information fusion and low-level vision technique, image fusion integrates complementary information from source images to yield an informative fused image. A few attempts have been made in recent years to jointly realize image fusion and super-resolution. However, in real-world applications such as military reconnaissance and long-range detection missions, the target and background structures in multimodal images are easily corrupted, with low resolution and weak semantic information, which leads to suboptimal results in current fusion techniques. In response, we propose FS-Diff, a semantic guidance and clarity-aware joint image fusion and super-resolution method. FS-Diff unifies image fusion and super-resolution as a conditional generation problem. It leverages semantic guidance from the proposed clarity sensing mechanism for adaptive low-resolution perception and cross-modal feature extraction. Specifically, we initialize the desired fused result as pure Gaussian noise and introduce the bidirectional feature Mamba to extract the global features of the multimodal images. Moreover, utilizing the source images and semantics as conditions, we implement a random iterative denoising process via a modified U-Net network. This network istrained for denoising at multiple noise levels to produce high-resolution fusion results with cross-modal features and abundant semantic information. We also construct a powerful aerial view multiscene (AVMS) benchmark covering 600 pairs of images. Extensive joint image fusion and super-resolution experiments on six public and our AVMS datasets demonstrated that FS-Diff outperforms the state-of-the-art methods at multiple magnifications and can recover richer details and semantics in the fused images. The code is available at https://github.com/XylonXu01/FS-Diff.

URL PDF HTML 收藏
1401.0166 2014-01-03 cs.CV cs.AI physics.med-ph 88%

Medical Image Fusion: A survey of the state of the art

A. P. James, B. V. Dasarathy

专题命中 通用Image Fusion :image fusion(title,abstract);medical image fusion(title,abstract);分类 cs.CV;information fusion(comments)

Comments Information Fusion, 2014

详情
英文摘要

Medical image fusion is the process of registering and combining multiple images from single or multiple imaging modalities to improve the imaging quality and reduce randomness and redundancy in order to increase the clinical applicability of medical images for diagnosis and assessment of medical problems. Multi-modal medical image fusion algorithms and devices have shown notable achievements in improving clinical accuracy of decisions based on medical images. This review article provides a factual listing of methods and summarizes the broad scientific challenges faced in the field of medical image fusion. We characterize the medical image fusion research based on (1) the widely used image fusion methods, (2) imaging modalities, and (3) imaging of organs that are under study. This review concludes that even though there exists several open ended technological and scientific challenges, the fusion of medical images has proved to be useful for advancing the clinical reliability of using medical imaging for medical diagnostics and analysis, and is a scientific discipline that has the potential to significantly grow in the coming years.

URL PDF HTML 收藏
2005.08448 2020-05-19 eess.IV cs.CV cs.MM 88%

Deep Convolutional Sparse Coding Networks for Image Fusion

Shuang Xu, Zixiang Zhao, Yicheng Wang, Chunxia Zhang, Junmin Liu, Jiangshe Zhang

专题命中 通用Image Fusion :image fusion(title,abstract);multi-modal image fusion(abstract);infrared and visible(abstract);multi-exposure(abstract)

详情
英文摘要

Image fusion is a significant problem in many fields including digital photography, computational imaging and remote sensing, to name but a few. Recently, deep learning has emerged as an important tool for image fusion. This paper presents three deep convolutional sparse coding (CSC) networks for three kinds of image fusion tasks (i.e., infrared and visible image fusion, multi-exposure image fusion, and multi-modal image fusion). The CSC model and the iterative shrinkage and thresholding algorithm are generalized into dictionary convolution units. As a result, all hyper-parameters are learned from data. Our extensive experiments and comprehensive comparisons reveal the superiority of the proposed networks with regard to quantitative evaluation and visual inspection.

URL PDF HTML 收藏
2606.12303 2026-06-11 cs.CV 新提交 88%

From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion

从二维网格到一维标记:重塑多模态图像融合的共享表示

Yuchen Xian, Yunqiu Xu, Yang He, Yi Yang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 通用Image Fusion :image fusion(title,abstract);multimodal image fusion(title,abstract);分类 cs.CV

AI总结 提出基于冻结预训练图像标记器的紧凑一维标记接口,通过选择性标记编辑(STE)稀疏更新关键标记,在保持融合骨干网络不变的同时引导全局外观一致性,实现全局连贯与局部保真的最佳平衡。

Comments Accepted at the 43rd International Conference on Machine Learning (ICML 2026)

详情
AI中文摘要

多模态图像融合旨在将来自不同模态的互补信息整合到融合图像中,该图像在保持全局一致外观的同时保留丰富的局部细节。现有方法在二维特征网格上构建共享表示,这些表示擅长建模局部结构,但对图像级全局外观因素的利用有限。为平衡这些目标,我们引入了一种基于冻结预训练图像标记器的紧凑一维标记接口,用于建模非局部外观/基因素。我们的设计不是将标记器用作重建骨干,而是将一维标记空间用作全局载体,同时保留用于局部结构恢复的二维空间路径。具体来说,我们引入了选择性标记编辑(STE),它稀疏地更新/替换一小部分关键标记,提供了一种轻量级机制来引导全局外观一致性,同时保持融合骨干网络不变并避免额外损失。在四个常用基准上的实验表明,我们的方法实现了最佳整体性能,在全局连贯性和局部保真度方面均具有一致的多指标改进。项目页面:此 https URL

英文摘要

Multimodal image fusion aims to integrate complementary information from different modalities into a fused image that preserves rich local details while maintaining globally consistent appearance. Existing approaches build shared representations on 2D feature grids, which excel at modeling local structures but offer limited leverage over image-level global appearance factors. To balance these objectives, we introduce a compact 1D token interface based on a frozen pretrained image tokenizer for modeling non-local appearance/base factors. Rather than using the tokenizer as a reconstruction backbone, our design uses the 1D token space as a global carrier while retaining the 2D spatial pathway for local structure restoration. Specifically, we introduce Selective Token Editing (STE), which sparsely updates/replaces a small set of critical tokens, providing a lightweight mechanism to steer global appearance coherence while keeping the fusion backbone unchanged and avoiding extra losses. Experiments on four commonly used benchmarks show that our method achieves the best overall performance, with consistent, multi-metric improvements in both global coherence and local fidelity. Project page: https://zju-xyc.github.io/1D-Fusion-Project-Page/

URL PDF HTML 收藏
2603.21129 2026-03-24 cs.CV 88%

ReDiffuse: Rotation Equivariant Diffusion Model for Multi-focus Image Fusion

ReDiffuse:用于多焦点图像融合的旋转等价扩散模型

Bo Li, Tingting Bao, Lingling Zhang, Weiping Fu, Yaxian Wang, Jun Liu

机构 * School of Computer Science and Technology, Xi'an Jiaotong University(西安交通大学计算机科学与技术学院) School of Information Engineering, Chang'an University(长安大学信息工程学院)

专题命中 通用Image Fusion :image fusion(title,abstract);multi-focus(title,abstract);分类 cs.CV

AI总结 本文提出ReDiffuse模型,通过在扩散网络中嵌入旋转等价性,解决多焦点图像融合中因模糊导致的结构失真问题,提升融合图像的结构一致性。

Comments 10 pages, 9 figures

详情
AI中文摘要

扩散模型在多焦点图像融合(MFIF)中表现出色,但其应用面临挑战,因为模糊会使对称几何结构变形,导致融合图像出现意外伪影。为此,本文提出ReDiffuse,通过构建端到端的旋转等价扩散模型,确保融合结果忠实保留输入图像的原始方向和结构一致性。通过四个数据集的全面评估,ReDiffuse在六个评估指标上实现了2.8%-6.64%的提升。

英文摘要

Diffusion models have achieved impressive performance on multi-focus image fusion (MFIF). However, a key challenge in applying diffusion models to the ill-posed MFIF problem is that defocus blur can make common symmetric geometric structures (e.g., textures and edges) appear warped and deformed, often leading to unexpected artifacts in the fused images. Therefore, embedding rotation equivariance into diffusion networks is essential, as it enables the fusion results to faithfully preserve the original orientation and structural consistency of geometric patterns underlying the input images. Motivated by this, we propose ReDiffuse, a rotation-equivariant diffusion model for MFIF. Specifically, we carefully construct the basic diffusion architectures to achieve end-to-end rotation equivariance. We also provide a rigorous theoretical analysis to evaluate its intrinsic equivariance error, demonstrating the validity of embedding equivariance structures. ReDiffuse is comprehensively evaluated against various MFIF methods across four datasets (Lytro, MFFW, MFI-WHU, and Road-MF). Results demonstrate that ReDiffuse achieves competitive performance, with improvements of 0.28-6.64\% across six evaluation metrics. The code is available at https://github.com/MorvanLi/ReDiffuse.

URL PDF HTML 收藏
2509.17704 2026-03-16 cs.CV 88%

Neurodynamics-Driven Coupled Neural P Systems for Multi-Focus Image Fusion

基于神经动力学的耦合神经P系统用于多焦点图像融合

Bo Li, Yunkuo Lei, Tingting Bao, Hang Yan, Yaxian Wang, Weiping Fu, Lingling Zhang, Jun Liu

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) Ministry of Education Key Laboratory of Intelligent Networks and Network Security, China(教育部智能网络与网络安全重点实验室) Shaanxi Province Key Laboratory of Big Data Knowledge Engineering(陕西省大数据知识工程重点实验室) Chang’an University(长安大学)

专题命中 通用Image Fusion :image fusion(title,abstract);multi-focus(title,abstract);分类 cs.CV

AI总结 本文提出基于神经动力学的耦合神经P系统,通过分析神经动力学约束生成高质量决策图,提升多焦点图像融合的准确性。

Comments Accepted by CVPR2026

详情
AI中文摘要

多焦点图像融合(MFIF)是图像处理中的关键技术,其核心挑战是生成具有精确边界决策图。传统基于启发式规则和深度学习的方法难以生成高质量决策图。为此,本文引入基于脉冲机制的第三代神经计算模型——神经动力学驱动的耦合神经P(CNP)系统,深入分析模型的神经动力学以确定网络参数与输入信号之间的约束关系,从而避免神经元异常连续放电,准确区分聚焦与非聚焦区域,生成高质量决策图。基于此分析,提出针对挑战性MFIF任务的神经动力学驱动CNP融合模型(ND-CNPFuse)。不同于现有决策图生成方法,ND-CNPFuse通过将源图像映射到可解释的脉冲矩阵中区分聚焦与非聚焦区域,通过比较脉冲数量直接生成准确决策图,无需后续处理。实验结果表明,ND-CNPFuse在四个经典MFIF数据集(Lytro、MFFW、MFI-WHU和Real-MFF)上实现了新的SOTA性能。代码可在https://github.com/MorvanLi/ND-CNPFuse获取。

英文摘要

Multi-focus image fusion (MFIF) is a crucial technique in image processing, with a key challenge being the generation of decision maps with precise boundaries. However, traditional methods based on heuristic rules and deep learning methods with black-box mechanisms are difficult to generate high-quality decision maps. To overcome this challenge, we introduce neurodynamics-driven coupled neural P (CNP) systems, which are third-generation neural computation models inspired by spiking mechanisms, to enhance the accuracy of decision maps. Specifically, we first conduct an in-depth analysis of the model's neurodynamics to identify the constraints between the network parameters and the input signals. This solid analysis avoids abnormal continuous firing of neurons and ensures the model accurately distinguishes between focused and unfocused regions, generating high-quality decision maps for MFIF. Based on this analysis, we propose a Neurodynamics-Driven CNP Fusion model (ND-CNPFuse) tailored for the challenging MFIF task. Unlike current ideas of decision map generation, ND-CNPFuse distinguishes between focused and unfocused regions by mapping the source image into interpretable spike matrices. By comparing the number of spikes, an accurate decision map can be generated directly without any post-processing. Extensive experimental results show that ND-CNPFuse achieves new state-of-the-art performance on four classical MFIF datasets, including Lytro, MFFW, MFI-WHU, and Real-MFF. The code is available at https://github.com/MorvanLi/ND-CNPFuse.

URL PDF HTML 收藏
2603.07120 2026-03-10 cs.CV 88%

Inter-Image Pixel Shuffling for Multi-focus Image Fusion

跨图像像素洗牌用于多焦点图像融合

Huangxing Lin, Rongrong Ma, Cheng Wang

机构 * College of Computer Science and Technology, Huaqiao University(华交大学计算机科学与技术学院)

专题命中 通用Image Fusion :image fusion(title,abstract);multi-focus(title,abstract);分类 cs.CV

AI总结 本文提出IPS方法,通过像素洗牌实现多焦点图像融合,无需实际多焦点图像即可学习融合过程,提升融合质量。

详情
AI中文摘要

多焦点图像融合旨在将多个部分聚焦图像合并为一个全聚焦图像。尽管深度学习在这一任务中展现出潜力,但其效果常受到合适训练数据稀缺的限制。本文介绍了Inter-image Pixel Shuffling(IPS),一种新颖的方法,使神经网络能够在不需实际多焦点图像的情况下学习多焦点图像融合。IPS将任务重新表述为像素级分类问题,目标是在每个空间位置的像素组中识别出聚焦的像素。在该方法中,清晰光学图像的像素被视为聚焦,而相同图像的低通滤波版本的像素被视为模糊。通过在原始和滤波图像中相同空间位置的聚焦和模糊像素随机洗牌,IPS生成的训练数据在保持空间结构的同时混合了聚焦-模糊信息。模型被训练以从每个空间对齐的像素组中选择聚焦像素,从而通过聚合输入中的锐利内容来学习重建全聚焦图像。为进一步提高融合质量,IPS采用跨图像融合网络,将卷积神经网络的局部表示能力与状态空间模型的长距离建模能力相结合。这种设计有效利用了空间细节和上下文信息,以产生高质量的融合结果。实验结果表明,IPS在不训练于多焦点图像的情况下显著优于现有多焦点图像融合方法。

英文摘要

Multi-focus image fusion aims to combine multiple partially focused images into a single all-in-focus image. Although deep learning has shown promise in this task, its effectiveness is often limited by the scarcity of suitable training data. This paper introduces Inter-image Pixel Shuffling (IPS), a novel method that allows neural networks to learn multi-focus image fusion without requiring actual multi-focus images. IPS reformulates the task as a pixel-wise classification problem, where the goal is to identify the focused pixel from a pixel group at each spatial position. In this method, pixels from a clear optical image are treated as focused, while pixels from a low-pass filtered version of the same image are considered defocused. By randomly shuffling the focused and defocused pixels at identical spatial positions in the original and filtered images, IPS generates training data that preserves spatial structure while mixing focus-defocus information. The model is trained to select the focused pixel from each spatially aligned pixel group, thus learning to reconstruct an all-in-focus image by aggregating sharp content from the input. To further enhance fusion quality, IPS adopts a cross-image fusion network that integrates the localized representation power of convolutional neural networks with the long-range modeling capabilities of state space models. This design effectively leverages both spatial detail and contextual information to produce high-quality fused results. Experimental results indicate that IPS significantly outperforms existing multi-focus image fusion methods, even without training on multi-focus images.

URL PDF HTML 收藏
2409.01728 2026-03-02 cs.CV 88%

Shuffle Mamba: State Space Models with Random Shuffle for Multi-Modal Image Fusion

Shuffle Mamba:基于随机洗牌的态空间模型用于多模态图像融合

Ke Cao, Xuanhua He, Tao Hu, Chengjun Xie, Man Zhou, Jie Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Intelligent Machines(智能机器研究所) Hefei Institutes of Physical Science, Chinese Academy of Sciences(中国科学院合肥物质科学研究院) Intelligent Agriculture Engineering Laboratory of Anhui Province, Institute of Intelligent Machines(安徽省智能农业工程实验室,智能机器研究所)

专题命中 通用Image Fusion :image fusion(title,abstract);multi-modal image fusion(title,abstract);分类 cs.CV

AI总结 Shuffle Mamba通过引入随机洗牌策略和逆洗牌,解决多模态图像融合中固定扫描策略带来的偏见问题,提升融合质量。

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology

详情
AI中文摘要

多模态图像融合通过整合不同模态的互补信息来生成增强且信息丰富的图像。尽管状态空间模型,如Mamba,在长距离建模方面具有线性复杂度的优势,但大多数基于Mamba的方法使用固定扫描策略,这可能会引入有偏的先验信息。为缓解这一问题,我们提出了一种新的贝叶斯启发式的扫描策略,称为随机洗牌,并辅以理论上可行的逆洗牌,以保持信息协调不变性,旨在消除与固定序列扫描相关的偏见。基于这一转换对,我们定制了Shuffle Mamba框架,穿透模态感知的信息表示和跨模态信息交互,通过空间和通道轴确保稳健的交互和无偏的全局感受野,以实现多模态图像融合。此外,我们开发了一种基于蒙特卡洛平均的测试方法,以确保模型的输出更接近预期结果。在多个多模态图像融合任务上的广泛实验表明,我们提出的方法的有效性,相较于最先进的替代方法,产生了出色的融合质量。代码可在https://github.com/caoke-963/Shuffle-Mamba上获得。

英文摘要

Multi-modal image fusion integrates complementary information from different modalities to produce enhanced and informative images. Although State-Space Models, such as Mamba, are proficient in long-range modeling with linear complexity, most Mamba-based approaches use fixed scanning strategies, which can introduce biased prior information. To mitigate this issue, we propose a novel Bayesian-inspired scanning strategy called Random Shuffle, supplemented by a theoretically feasible inverse shuffle to maintain information coordination invariance, aiming to eliminate biases associated with fixed sequence scanning. Based on this transformation pair, we customized the Shuffle Mamba Framework, penetrating modality-aware information representation and cross-modality information interaction across spatial and channel axes to ensure robust interaction and an unbiased global receptive field for multi-modal image fusion. Furthermore, we develop a testing methodology based on Monte-Carlo averaging to ensure the model's output aligns more closely with expected results. Extensive experiments across multiple multi-modal image fusion tasks demonstrate the effectiveness of our proposed method, yielding excellent fusion quality compared to state-of-the-art alternatives. The code is available at https://github.com/caoke-963/Shuffle-Mamba.

URL PDF HTML 收藏
2601.05538 2026-01-12 cs.CV 88%

DIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multi-Modal Image Fusion

DIFF-MF: 一种基于差分驱动的通道-空间状态空间模型用于多模态图像融合

Yiming Sun, Zifan Ye, Qinghua Hu, Pengfei Zhu

机构 * School of Automation, Southeast University(东南大学自动化学院) Low-Altitude Intelligence Lab, Xiong’an National Innovation Center Technology Co., Ltd.(雄安国家创新中心技术有限公司低空智能实验室) Xiong’an Guochuang Lantian Technology Co., Ltd.(雄安国创蓝天科技有限公司) School of Artificial Intelligence, Tianjin University(天津大学人工智能学院)

专题命中 通用Image Fusion :image fusion(title,abstract);multi-modal image fusion(title,abstract);分类 cs.CV

AI总结 DIFF-MF通过差分驱动的通道-空间状态空间模型,有效整合多模态图像信息,提升融合图像的质量和显著性。

详情
AI中文摘要

多模态图像融合旨在整合多个源图像中的互补信息,以生成高质量且内容丰富的融合图像。尽管现有基于状态空间模型的方法在高计算效率下取得了满意的效果,但它们往往倾向于以牺牲可见细节为代价过度优先考虑红外强度,或相反,保留可见结构的同时削弱热目标的显著性。为克服这些挑战,我们提出了DIFF-MF,一种新的基于差分驱动的通道-空间状态空间模型用于多模态图像融合。我们的方法利用模态之间的特征差异图来指导特征提取,随后在通道和空间维度上进行融合过程。在通道维度上,通道交换模块通过跨注意力双状态空间建模增强通道间的交互,实现自适应特征重加权。在空间维度上,空间交换模块采用跨模态状态空间扫描以实现全面的空间融合。通过高效捕捉全局依赖性同时保持线性计算复杂度,DIFF-MF有效整合了互补的多模态特征。在驾驶场景和低空无人机数据集上的实验结果表明,我们的方法在视觉质量和定量评估上均优于现有方法。

英文摘要

Multi-modal image fusion aims to integrate complementary information from multiple source images to produce high-quality fused images with enriched content. Although existing approaches based on state space model have achieved satisfied performance with high computational efficiency, they tend to either over-prioritize infrared intensity at the cost of visible details, or conversely, preserve visible structure while diminishing thermal target salience. To overcome these challenges, we propose DIFF-MF, a novel difference-driven channel-spatial state space model for multi-modal image fusion. Our approach leverages feature discrepancy maps between modalities to guide feature extraction, followed by a fusion process across both channel and spatial dimensions. In the channel dimension, a channel-exchange module enhances channel-wise interaction through cross-attention dual state space modeling, enabling adaptive feature reweighting. In the spatial dimension, a spatial-exchange module employs cross-modal state space scanning to achieve comprehensive spatial fusion. By efficiently capturing global dependencies while maintaining linear computational complexity, DIFF-MF effectively integrates complementary multi-modal features. Experimental results on the driving scenarios and low-altitude UAV datasets demonstrate that our method outperforms existing approaches in both visual quality and quantitative evaluation.

URL PDF HTML 收藏