arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

Image Fusion

围绕 image fusion 的图像融合方法、数据集、评测与应用。

至 收录 676
2510.12260 2025-10-15 cs.CV cs.LG eess.IV 79%

AngularFuse: A Closer Look at Angle-based Perception for Spatial-Sensitive Multi-Modality Image Fusion

Xiaopeng Liu, Yupei Lin, Sen Zhang, Xiao Wang, Yukai Shi, Liang Lin

机构 * School of Information Engineering, Guangdong University of Technology(广东技术大学信息工程学院) TikTok, ByteDance Inc(字节跳动) School of Computer Science, Anhui University(安徽大学计算机科学学院) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

专题命中 Image Fusion :image fusion(title,abstract)

Comments For the first time, angle-based perception was introduced into the multi-modality image fusion task

详情
英文摘要

Visible-infrared image fusion is crucial in key applications such as autonomous driving and nighttime surveillance. Its main goal is to integrate multimodal information to produce enhanced images that are better suited for downstream tasks. Although deep learning based fusion methods have made significant progress, mainstream unsupervised approaches still face serious challenges in practical applications. Existing methods mostly rely on manually designed loss functions to guide the fusion process. However, these loss functions have obvious limitations. On one hand, the reference images constructed by existing methods often lack details and have uneven brightness. On the other hand, the widely used gradient losses focus only on gradient magnitude. To address these challenges, this paper proposes an angle-based perception framework for spatial-sensitive image fusion (AngularFuse). At first, we design a cross-modal complementary mask module to force the network to learn complementary information between modalities. Then, a fine-grained reference image synthesis strategy is introduced. By combining Laplacian edge enhancement with adaptive histogram equalization, reference images with richer details and more balanced brightness are generated. Last but not least, we introduce an angle-aware loss, which for the first time constrains both gradient magnitude and direction simultaneously in the gradient domain. AngularFuse ensures that the fused images preserve both texture intensity and correct edge orientation. Comprehensive experiments on the MSRS, RoadScene, and M3FD public datasets show that AngularFuse outperforms existing mainstream methods with clear margin. Visual comparisons further confirm that our method produces sharper and more detailed results in challenging scenes, demonstrating superior fusion capability.

URL PDF HTML 收藏
2502.14493 2025-02-21 cs.CV cs.LG 79%

CrossFuse: Learning Infrared and Visible Image Fusion by Cross-Sensor Top-K Vision Alignment and Beyond

Yukai Shi, Cidan Shi, Zhipeng Weng, Yin Tian, Xiaoyu Xian, Liang Lin

专题命中 Image Fusion :image fusion(title,abstract)

Comments IEEE T-CSVT. We mainly discuss the out-of-distribution challenges in infrared and visible image fusion

详情
英文摘要

Infrared and visible image fusion (IVIF) is increasingly applied in critical fields such as video surveillance and autonomous driving systems. Significant progress has been made in deep learning-based fusion methods. However, these models frequently encounter out-of-distribution (OOD) scenes in real-world applications, which severely impact their performance and reliability. Therefore, addressing the challenge of OOD data is crucial for the safe deployment of these models in open-world environments. Unlike existing research, our focus is on the challenges posed by OOD data in real-world applications and on enhancing the robustness and generalization of models. In this paper, we propose an infrared-visible fusion framework based on Multi-View Augmentation. For external data augmentation, Top-k Selective Vision Alignment is employed to mitigate distribution shifts between datasets by performing RGB-wise transformations on visible images. This strategy effectively introduces augmented samples, enhancing the adaptability of the model to complex real-world scenarios. Additionally, for internal data augmentation, self-supervised learning is established using Weak-Aggressive Augmentation. This enables the model to learn more robust and general feature representations during the fusion process, thereby improving robustness and generalization. Extensive experiments demonstrate that the proposed method exhibits superior performance and robustness across various conditions and environments. Our approach significantly enhances the reliability and stability of IVIF tasks in practical applications.

URL PDF HTML 收藏
2103.00940 2021-08-04 eess.IV cs.CV 79%

LADMM-Net: An Unrolled Deep Network For Spectral Image Fusion From Compressive Data

Juan Marcos Ramírez, José Ignacio Martínez Torre, Henry Arguello Fuentes

专题命中 Image Fusion :image fusion(title,abstract)

Comments 29 pages, 15 figures, 4 tables

Journal ref Juan Marcos Ramirez, Jose Ignacio Martinez-Torre, and Henry Arguello, "LADMM-Net: An Unrolled Deep Network For Spectral Image Fusion From Compressive Data", Signal Processing, vol. 189, Dec 2021, 108239

详情
英文摘要

Image fusion aims at estimating a high-resolution spectral image from a low-spatial-resolution hyperspectral image and a low-spectral-resolution multispectral image. In this regard, compressive spectral imaging (CSI) has emerged as an acquisition framework that captures the relevant information of spectral images using a reduced number of measurements. Recently, various image fusion methods from CSI measurements have been proposed. However, these methods exhibit high running times and face the challenging task of choosing sparsity-inducing bases. In this paper, a deep network under the algorithm unrolling approach is proposed for fusing spectral images from compressive measurements. This architecture, dubbed LADMM-Net, casts each iteration of a linearized version of the alternating direction method of multipliers into a processing layer whose concatenation deploys a deep network. The linearized approach enables obtaining fusion estimates without resorting to costly matrix inversions. Furthermore, this approach exploits the benefits of learnable transforms to estimate the image details included in both the auxiliary variable and the Lagrange multiplier. Finally, the performance of the proposed technique is evaluated on two spectral image databases and one dataset captured at the laboratory. Extensive simulations show that the proposed method outperforms the state-of-the-art approaches that fuse spectral images from compressive measurements.

URL PDF HTML 收藏
1307.2440 2013-07-10 cs.CV 79%

Image Fusion Technologies In Commercial Remote Sensing Packages

Firouz Abdullah Al-Wassai, N. V. Kalyankar

专题命中 Image Fusion :image fusion(title,abstract)

Comments Keywords: Commercial Processing Systems, Image Fusion, quality evaluation

Journal ref Journal of Global Research in Computer Science, 4 (5), May 2013, 44-50

详情
英文摘要

Several remote sensing software packages are used to the explicit purpose of analyzing and visualizing remotely sensed data, with the developing of remote sensing sensor technologies from last ten years. Accord-ing to literature, the remote sensing is still the lack of software tools for effective information extraction from remote sensing data. So, this paper provides a state-of-art of multi-sensor image fusion technologies as well as review on the quality evaluation of the single image or fused images in the commercial remote sensing pack-ages. It also introduces program (ALwassaiProcess) developed for image fusion and classification.

URL PDF HTML 收藏
1107.4396 2012-09-18 cs.CV 79%

The IHS Transformations Based Image Fusion

Firouz Abdullah Al-Wassai, N. V. Kalyankar, Ali A. Al-Zuky

专题命中 Image Fusion :image fusion(title,abstract)

Comments Image Fusion, Color Models, IHS, HSV, HSL, YIQ, transformations

Journal ref Journal-ref: International Journal of Advanced Research in Computer Science,Volume 2, No. 5, Sept-Oct 2011,www.ijarcs.info

详情
英文摘要

The IHS sharpening technique is one of the most commonly used techniques for sharpening. Different transformations have been developed to transfer a color image from the RGB space to the IHS space. Through literature, it appears that, various scientists proposed alternative IHS transformations and many papers have reported good results whereas others show bad ones as will as not those obtained which the formula of IHS transformation were used. In addition to that, many papers show different formulas of transformation matrix such as IHS transformation. This leads to confusion what is the exact formula of the IHS transformation?. Therefore, the main purpose of this work is to explore different IHS transformation techniques and experiment it as IHS based image fusion. The image fusion performance was evaluated, in this study, using various methods to estimate the quality and degree of information improvement of a fused image quantitatively.

URL PDF HTML 收藏
1107.3348 2012-09-18 cs.CV 79%

Arithmetic and Frequency Filtering Methods of Pixel-Based Image Fusion Techniques

Firouz Abdullah Al-Wassai, N. V. Kalyankar, Ali A. Al-Zuky

专题命中 Image Fusion :image fusion(title,abstract)

Comments Image Fusion, Pixel-Based Fusion, Brovey Transform, Color Normalized, High-Pass Filter, Modulation, Wavelet transform

Journal ref Journal-ref: International Journal of Advanced Research in Computer Science,Volume 2, No. 5, Sept-Oct 2011,www.ijarcs.info

详情
英文摘要

In remote sensing, image fusion technique is a useful tool used to fuse high spatial resolution panchromatic images (PAN) with lower spatial resolution multispectral images (MS) to create a high spatial resolution multispectral of image fusion (F) while preserving the spectral information in the multispectral image (MS).There are many PAN sharpening techniques or Pixel-Based image fusion techniques that have been developed to try to enhance the spatial resolution and the spectral property preservation of the MS. This paper attempts to undertake the study of image fusion, by using two types of pixel-based image fusion techniques i.e. Arithmetic Combination and Frequency Filtering Methods of Pixel-Based Image Fusion Techniques. The first type includes Brovey Transform (BT), Color Normalized Transformation (CN) and Multiplicative Method (MLT). The second type include High-Pass Filter Additive Method (HPFA), High-Frequency-Addition Method (HFA) High Frequency Modulation Method (HFM) and The Wavelet transform-based fusion method (WT). This paper also devotes to concentrate on the analytical techniques for evaluating the quality of image fusion (F) by using various methods including Standard Deviation (SD), Entropy(En), Correlation Coefficient (CC), Signal-to Noise Ratio (SNR), Normalization Root Mean Square Error (NRMSE) and Deviation Index (DI) to estimate the quality and degree of information improvement of a fused image quantitatively.

URL PDF HTML 收藏
2607.03643 2026-07-07 eess.IV cs.CV 新提交 78%

Model Confidence-Guided Multi-Image Fusion of Fundus Images for Diabetic Retinopathy Diagnosis

模型置信度引导的多图像融合眼底图像糖尿病视网膜病变诊断方法

Ananya Raghu, Anisha Raghu, Alice S. Tang, Yannis M. Paulus, Tyson N. Kim, Tomiko T. Oskotsky

机构 * Massachusetts Institute of Technology(麻省理工学院) Wilmer Eye Institute, Department of Ophthalmology, Johns Hopkins University(约翰霍普金斯大学威尔默眼科研究所) Department of Biomedical Engineering, Johns Hopkins University(约翰霍普金斯大学生物医学工程系) Bakar Computational Health Sciences Institute, University of California San Francisco(加州大学旧金山分校巴卡计算健康科学研究所) Department of Ophthalmology, University of California San Francisco(加州大学旧金山分校眼科系) Division of Clinical Informatics and Digital Transformation, University of California San Francisco(加州大学旧金山分校临床信息学与数字化转型 division)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 针对资源有限地区糖网病早期筛查需求,提出置信度引导的多图像融合框架,整合多眼底视图提升诊断置信度与准确率,性能优于传统图像质量级联流水线,适配低延迟移动筛查场景。

详情
AI中文摘要

目的:在就医渠道有限的中低收入国家,眼部疾病的早期筛查至关重要。本研究探讨置信度引导的多图像糖尿病视网膜病变诊断框架能否将图像过滤与置信度感知预测相结合,在图像采集端实现可靠筛查。方法:开发一种多图像融合方法,聚合多视角眼底图像以提升诊断置信度与平衡准确率,该方法利用置信度识别不可靠预测,必要时提示重新拍摄。对比三类方案:(1)每患者单张图像的级联图像质量与疾病诊断流水线,(2)基于置信度的预测方案,(3)本文提出的基于置信度的多图像融合流水线。所有方法均基于RETFoundGreen骨干网络,在mBRSET(n=1234)和BRSET(n=7599)数据集上完成评估。结果:在70%覆盖率下,本文方法在mBRSET上达到91%的平衡准确率,在BRSET上达到97%的平衡准确率,较级联过滤方案分别提升约12%和6%。图像质量级联方案在mBRSET上的灵敏度为61%,在BRSET上为86%,而本文框架在50%覆盖率下的灵敏度分别可达94%和96%。结论:人工标注的质量标签与诊断性能的相关性较弱,基于置信度的过滤方案始终优于基于图像质量的级联流水线。转化意义:采用基于置信度的多图像融合技术可为患者提供更可靠的预测,减少筛查过程中的误诊情况。该框架采用轻量骨干网络,且每张图像仅需单次推理,可适配资源受限场景下的低延迟移动筛查系统。

英文摘要

Purpose: Early screening for eye diseases is critical in low- and middle-income countries where access to care is limited. We investigate whether a confidence-guided, multi-image diabetic retinopathy diagnosis framework can integrate image filtering with confidence-aware predictions for reliable screening at capture. Methods: We develop a multi-image fusion method that aggregates retinal views to improve confidence and balanced accuracy. Our method uses confidence to identify unreliable predictions, prompting retakes when needed. We compare: (1) a cascaded image-quality and disease diagnosis pipeline using a single image per patient, (2) confidence-based prediction, and (3) our confidence-based multi-image fusion pipeline. All methods are evaluated using a RETFoundGreen backbone on the mBRSET (n = 1,234) and BRSET (n = 7,599) datasets. Results: At 70% coverage, our method achieves 91% balanced accuracy on mBRSET and 97% on BRSET, improvements of ~12% and ~6%, respectively, over cascade filtering. The image-quality cascade reaches sensitivities of 61% on mBRSET and 86% on BRSET, whereas our framework reaches 94% and 96%, respectively, at 50% coverage. Conclusions: Human-annotated quality labels are weakly associated with diagnostic performance, and confidence-based filtering consistently outperforms image quality-based cascaded pipelines. Translational Relevance: Using confidence-based multi-image fusion, patients receive more reliable predictions, reducing incorrect diagnoses during screening. The lightweight backbone and single inference pass per image make the framework compatible with low-latency mobile screening systems in resource-limited settings.

URL PDF HTML 收藏
2607.02572 2026-07-07 cs.CV cs.AI 新提交 78%

Additive Causal Construction for Transferable and Reconfigurable Cross-System Learning in Multi-Source Image Fusion

多源图像融合中用于可转移和可重构跨系统学习的加性因果构建

Zhizhong Fu, Wei Zhou, Zhaoyang Jiang, Yulong Lin, Yifu Hou, Xiaorong Ding, Qiang Yan, Yifan Chen

机构 * School of Life Science and Technology, University of Electronic Science and Technology of China(电子科技大学生命科学与技术学院) School of Health and Wellbeing, University of Glasgow(格拉斯哥大学健康与幸福学院) Department of Organ Transplantation, Sichuan Provincial People’s Hospital, University of Electronic Science and Technology of China(电子科技大学附属四川省人民医院器官移植科) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) Department of Radiology, Huzhou Maternity & Child Health Care Hospital(湖州市妇幼保健院放射科) Hepatological Surgery Department, Huzhou Central Hospital, Fifth School of Clinical Medicine of Zhejiang Chinese Medical University(浙江中医药大学附属湖州中医院第五临床医学院湖州市中心医院肝外科)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 针对多源图像融合中跨系统差异和纠缠问题,提出加性因果构建框架,通过干预一致性建立共享因果‘锚’实现因果图可转移,将融合过程形式化为因果构建并量化不确定性确保可重构,改进因果表示学习。

详情
AI中文摘要

在多源图像融合场景中,异构输入由不同生成机制驱动,可视为多个因果系统的组合。融合过程中常出现跨系统差异(CSD)和跨系统纠缠(CSE),导致分布外(OOD)预测性能显著下降。为解决这些问题,我们提出加性因果构建(ACC)框架……

英文摘要

In multi-source image fusion scenarios, heterogeneous inputs are typically driven by distinct generative mechanisms and can be viewed as a composition of multiple causal systems. However, cross-system discrepancy (CSD) and cross-system entanglement (CSE) commonly arise during the fusion process, often leading to significant performance degradation under out-of-distribution (OOD) predictions. To address the CSD and CSE issues, we propose the additive causal construction (ACC) framework, which characterizes information fusion at two levels: firstly, it establishes causal "anchors" shared among multiple systems through intervention consistency to enable causal graph transferability (CGT); and secondly, it formalizes the fusion process as causal construction and models the reliability of constructed paths through uncertainty quantification to ensure causal graph reconfigurability (CGR). Building upon this, we revisit the traditional causal representation learning (CRL) with ACC and propose ACC-CRL as a learnable instantiation of the framework. The method explores joint causal content representations across systems via content-mechanism decoupling, and performs response alignment under shared anchors to mitigate CSD. Furthermore, it incorporates structural uncertainty to adaptively regulate the fusion process, thereby suppressing unstable CSE. We conduct systematic experiments on synthetic data (ColorMNIST) and real-world multi-center medical imaging tasks (microvascular invasion (MVI) prediction). The results demonstrate that the proposed method significantly improves OOD generalization while maintaining in-distribution (ID) performance, validating the effectiveness and robustness of the ACC-CRL strategy based on mechanism alignment and uncertainty modeling in open environments.

URL PDF HTML 收藏
2603.16130 2026-07-03 cs.CV 版本更新 78%

EPOFusion: Exposure aware Progressive Optimization Method for Infrared and Visible Image Fusion

EPOFusion:一种针对红外和可见图像融合的曝光感知渐进优化方法

Zhiwei Wang, Defeng He, Li Zhao, Xiaoqin Zhang, Yuxing Li, Edmund Y. Lam

机构 * College of Information Engineering, Zhejiang University of Technology(浙江工业大学信息工程学院) College of Information Science and Technology, Zhejiang Shuren University(浙江树人大学信息科技学院) Department of Electrical and Electronic Engineering, The University of Hong Kong(香港大学电机电子工程系)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出EPOFusion方法,通过引入指导模块和迭代解码器提升曝光区域的融合效果,同时构建首个高质红外引导标注的曝光数据集,实验表明其在视觉保真度和下游任务表现上优于现有方法。

详情
AI中文摘要

在实际场景中,过曝常导致关键视觉信息丢失,现有红外与可见融合方法在高亮区域表现不佳。为此,我们提出EPOFusion,一种曝光感知融合模型。具体而言,引入指导模块以帮助编码器从过曝区域提取细粒度红外特征。同时,设计包含多尺度上下文融合模块的迭代解码器,逐步增强融合图像,确保细节一致性和高质量视觉效果。最后,采用自适应损失函数动态约束融合过程,在不同曝光条件下实现模态间的有效平衡。为增强曝光感知,我们构建了首个红外与可见过曝数据集(IVOE),包含高质量红外引导标注。大量实验表明,EPOFusion在过曝区域保持红外线索,同时在非过曝区域实现视觉忠实融合,从而提升视觉保真度和下游任务性能。代码、融合结果和IVOE数据集将在https://github.com/warren-wzw/EPOFusion.git发布。

英文摘要

Overexposure caused by strong daylight and oncoming headlights frequently overwhelms visible sensors, resulting in critical information loss in visual perception. Infrared and visible image fusion can compensate for such degradation via multimodal complementarity. However, most fusion methods lack region-aware optimization for overexposed areas and cannot effectively exploit infrared cues in saturated regions, resulting in insufficient infrared detail preservation or redundant information in the fused results. To address this, we propose EPOFusion, an exposure-aware fusion framework. It uses a spatial guidance module to selectively preserve informative infrared cues in overexposed regions. In addition, an iterative decoding head equipped with a multiscale context fusion module progressively refines fused representations, enabling effective infrared compensation in degraded regions while maintaining visual consistency in normal regions. The infrared and visible overexposure (IVOE) dataset is constructed with a synthetic training subset for controlled supervision and a real-world test subset for generalization assessment, supporting exposure-aware learning and evaluation. Extensive experiments on MSRS, FMB, and the proposed IVOE benchmark show that EPOFusion improves information preservation and visual fidelity, achieving an average full-image MI gain of 28.7% over the best competing methods. Qualitative results further demonstrate effective compensation in saturated regions, and downstream evaluations confirm its benefits under challenging overexposed conditions. Code, results, and the IVOE dataset will be made available at https://github.com/warren-wzw/EPOFusion.

URL PDF HTML 收藏
2606.31242 2026-07-01 cs.CV 新提交 78%

UHD-MFF: Shattering Barriers in Multi-Focus Ultra-High-Definition Image Fusion via Learnable Lookup Tables

UHD-MFF:通过可学习查找表打破多焦点超高清图像融合的障碍

Yibing Zhang, Xunpeng Yi, Qinglong Yan, Yeda Wang, Han Xu, Jiayi Ma

机构 * Electronic Information School, Wuhan University(武汉大学电子信息学院) School of Robotics, Wuhan University(武汉大学机器人学院) School of Automation, Southeast University(东南大学自动化学院)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 针对超高清多焦点图像融合中的数据、模型和部署三大障碍,提出首个大规模超高清数据集UHD-MFF和基于可学习查找表的UMF-LUT框架,实现实时4K融合。

Comments Accepted by ECCV 2026

详情
AI中文摘要

随着成像技术的进步,超高清图像在现代视觉应用中变得越来越重要。然而,现有的多焦点图像融合大多局限于低分辨率图像,在超高清场景下面临三大障碍:数据可用性、模型适应性和部署可行性,严重阻碍了其实际应用。为了打破这些障碍,首先,我们提出了UHD-MFF数据集,这是首个大规模超高清多焦点融合数据集。其次,我们提出了一种针对超高清图像的尺度专用查找表框架,称为UMF-LUT。它由粗区域查找表(C-LUT)和细节边缘查找表(D-LUT)组成。具体来说,C-LUT在低分辨率尺度上对多个梯度线索和语义线索进行联合查询,以实现区域级决策。同时,D-LUT在高分辨率尺度上运行,利用高效的拉普拉斯线索提供互补的边缘级决策信息。这种设计使得模型特别适合超高清多焦点图像融合。最后,它提供了强大的可部署性,计算开销极小,实现了实时4K多焦点融合,并在智能手机上显示出巨大潜力。大量实验表明,它在视觉保真度和定量指标上均优于最先进的方法。它有效地推动了多焦点图像融合向超高清成像场景的发展。代码可在以下网址获取:this https URL。

英文摘要

With the advancement of imaging technology, ultra-high-definition images have become increasingly essential in modern visual applications. However, existing multi-focus image fusion remains largely confined to low-resolution images and faces three major barriers in UHD scenarios, namely data availability, model adaptability, and deployment feasibility, which severely hinder its practical application. To shatter these barriers, first, we propose the UHD-MFF dataset, the first large-scale ultra-high-resolution multi-focus fusion dataset. Second, we propose a scale-specialized lookup-table framework tailored for ultra-high-resolution images, termed as UMF-LUT. It consists of Coarse-Region Lookup Table (C-LUT) and Detail-Edge Lookup Table (D-LUT). Specifically, C-LUT performs joint queries of multiple gradient cues and semantic cues at low-resolution scales to enable region-level decision-making. Also, D-LUT operates at high-resolution scales, leveraging efficient Laplacian cues to provide complementary edge-level decision information. Such a design makes the model particularly well-suited for ultra-high-resolution multi-focus image fusion. Finally, it offers strong deployability with minimal computational overhead, enabling real-time 4K multi-focus fusion and showing promising potential for smartphone. Extensive experiments demonstrate that it outperforms SOTA methods in both visual fidelity and quantitative metrics. It effectively advances the development of multi-focus image fusion toward ultra-high-resolution imaging scenarios. The code is available at https://github.com/zyb5/UHD-MFF.

URL PDF HTML 收藏
2606.26812 2026-06-26 cs.CV 新提交 78%

Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction

恶劣天气下的多模态图像融合:掩码引导的特征恢复与交互

Xilai Li, Xiaosong Li, Haishu Tan, Tao Ye, Huafeng Li, Hongbin Wang

机构 * Foshan University(佛山大学) China University of Mining and Technology, Beijing(中国矿业大学(北京)) Kunming University of Science and Technology(昆明理工大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 提出一种掩码引导的多模态图像融合方法,通过伪真实标签和掩码生成机制同时实现特征恢复与跨模态交互,在合成和真实数据集上超越现有方法。

Comments Accepted at ECCV 2026

详情
AI中文摘要

多模态图像融合(MMIF)通过利用不同模态的互补线索来增强场景表示。然而,恶劣天气会导致显著的图像退化,破坏特征表示,需要同时进行特征恢复和跨模态互补。现有方法在此类条件下往往难以进行有效的表示学习,限制了其实际性能。为了解决这些挑战,我们提出了一种掩码引导的MMIF方法,该方法整合了特征恢复和交互。我们首先引入“伪真实标签”以简化训练,促进更快更有效的特征学习。然后,我们基于融合结果与源图像之间的映射关系设计了一种掩码生成机制,量化了融合过程中每种模态的相对贡献。通过引入所提出的掩码引导的跨模态交叉注意力机制,网络被鼓励在模态交互期间选择性地关注信息特征,从而减轻了对“伪真实标签”静态分布过拟合的风险。此外,我们提出了一种掩码引导的学习策略和一种任务耦合的退化感知学习策略,以平衡特征恢复和交互。在合成和真实数据集上的大量实验表明,我们的方法在视觉质量、量化指标和下游任务方面均超越了现有最先进的方法。源代码可从此处获取:此 https URL。

英文摘要

Multi-modality image fusion (MMIF) enhances scene representation by exploiting complementary cues from different modalities. Adverse weather, however, causes significant image degradation, disrupting feature representation and requiring simultaneous feature restoration and cross-modal complementarity. Existing methods often struggle with effective representation learning under such conditions, limiting their practical performance. To address these challenges, we propose a mask-guided MMIF method that integrates feature restoration and interaction. We first introduce "Pseudo Ground Truth" to simplify training, promoting faster and more effective feature learning. Then, we design a mask generation mechanism based on the mapping relationship between the fused result and the source images, quantifying the relative contribution of each modality during the fusion process. By incorporating the proposed mask-guided cross-modal cross-attention mechanism, the network is encouraged to selectively attend to informative features during modality interaction, mitigating the risk of overfitting to the static distribution of the "Pseudo Ground Truth". Additionally, we propose a mask-guided learning strategy and a task-coupled degradation-aware learning strategy to balance feature restoration and interaction. Extensive experiments on synthetic and real-world datasets demonstrate that our method surpasses state-of-the-art approaches in visual quality, quantitative metrics, and downstream tasks. The source code is available at https://github.com/ixilai/AMG-Fuse.

URL PDF HTML 收藏
2606.12303 2026-06-11 cs.CV 新提交 78%

From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion

从二维网格到一维标记:重塑多模态图像融合的共享表示

Yuchen Xian, Yunqiu Xu, Yang He, Yi Yang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 提出基于冻结预训练图像标记器的紧凑一维标记接口,通过选择性标记编辑(STE)稀疏更新关键标记,在保持融合骨干网络不变的同时引导全局外观一致性,实现全局连贯与局部保真的最佳平衡。

Comments Accepted at the 43rd International Conference on Machine Learning (ICML 2026)

详情
AI中文摘要

多模态图像融合旨在将来自不同模态的互补信息整合到融合图像中,该图像在保持全局一致外观的同时保留丰富的局部细节。现有方法在二维特征网格上构建共享表示,这些表示擅长建模局部结构,但对图像级全局外观因素的利用有限。为平衡这些目标,我们引入了一种基于冻结预训练图像标记器的紧凑一维标记接口,用于建模非局部外观/基因素。我们的设计不是将标记器用作重建骨干,而是将一维标记空间用作全局载体,同时保留用于局部结构恢复的二维空间路径。具体来说,我们引入了选择性标记编辑(STE),它稀疏地更新/替换一小部分关键标记,提供了一种轻量级机制来引导全局外观一致性,同时保持融合骨干网络不变并避免额外损失。在四个常用基准上的实验表明,我们的方法实现了最佳整体性能,在全局连贯性和局部保真度方面均具有一致的多指标改进。项目页面:此 https URL

英文摘要

Multimodal image fusion aims to integrate complementary information from different modalities into a fused image that preserves rich local details while maintaining globally consistent appearance. Existing approaches build shared representations on 2D feature grids, which excel at modeling local structures but offer limited leverage over image-level global appearance factors. To balance these objectives, we introduce a compact 1D token interface based on a frozen pretrained image tokenizer for modeling non-local appearance/base factors. Rather than using the tokenizer as a reconstruction backbone, our design uses the 1D token space as a global carrier while retaining the 2D spatial pathway for local structure restoration. Specifically, we introduce Selective Token Editing (STE), which sparsely updates/replaces a small set of critical tokens, providing a lightweight mechanism to steer global appearance coherence while keeping the fusion backbone unchanged and avoiding extra losses. Experiments on four commonly used benchmarks show that our method achieves the best overall performance, with consistent, multi-metric improvements in both global coherence and local fidelity. Project page: https://zju-xyc.github.io/1D-Fusion-Project-Page/

URL PDF HTML 收藏
2606.07985 2026-06-09 cs.CV cs.CL 新提交 78%

FMRFusion: Frequency-Aware Multi-View Representation Learning for Heterogeneous Image Fusion

FMRFusion: 面向异质图像融合的频率感知多视图表示学习

Tao Zhoua, Yunlong Liu, Qinghui Chen, Zekai Zhang, Minlong Sun, Changlin Biana, Dagang Li, Wenmin Wang, Jinglin Zhang

机构 * Shandong University(山东大学) Macau University of Science and Technology(澳门科技大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 提出FMRFusion网络,通过多尺度结构感知模块、双线性频率分解和跨视图互补交互,结合流匹配优化,实现红外与可见光图像融合,在夜间场景表现优异。

详情
AI中文摘要

红外与可见光图像融合旨在生成保留重要目标信息和详细纹理的复合图像,整合两种异质模态。以往的图像融合方法通常采用单模块堆叠方式从两种模态中提取特征,然而这些方法可能导致对其独特特征的学习不完整,从而限制融合效果并在真实异质数据场景中降低鲁棒性。为解决这些问题,我们提出FMRFusion,一种用于异质图像融合的频率感知多视图表示学习网络。引入多尺度结构感知模块以有效捕捉判别性结构,提取细粒度局部结构和关键上下文信息。采用双线性频率分解机制将特征分离为高频和低频分量,实现不同频率域中局部细节和全局表示的联合建模。此外,融入跨视图互补交互以显式建模和融合反射光信息与辐射强度响应之间的互补特性,促进有效的跨视图交互。我们通过流匹配进一步改善融合结果的质量,通过学习从粗数据到高质量表示的变换逐步细化融合特征。在多个基准数据集上进行的大量实验表明,FMRFusion在一系列融合任务中实现了优越且一致的性能,尤其在夜间场景中表现突出。

英文摘要

Infrared and visible image fusion aims to generate a composite image that retains significant target information and preserves detailed textures, integrating two heterogeneous modalities. Previous image fusion methods typically adopt a single-module stacking approach to extract features from the two modalities. However, these approaches may result in incomplete learning of their distinct characteristics, thereby limiting the fusion effectiveness and constrain ing robustness in real-world heterogeneous data scenarios. To address these challenges, we propose FMRFusion, a frequency-aware multi-view representation learning network for Heterogeneous Image Fusion. A Multi-Scale Struc tural Perception Module is introduced to effectively capture discriminative structures, extracting fine-grained local structures and essential contextual information. A bilinear frequency decomposition mechanism is employed to sepa rate features into high-frequency and low-frequency components, enabling joint modeling of local details and global representations across different frequency domains. Moreover, a Cross-View Complementary Interaction is incorpo rated to explicitly model and fuse the complementary characteristics between reflected light information and radiative intensity responses, facilitating effective cross-view interaction. We further improve the Performance of the fused results by flow matching, which progressively refines the fused features by learning the transformation from coarse data to high-quality representations. Extensive experiments conducted on multiple benchmark datasets demonstrate that FMRFusion achieves superior and consistent performance across a range of fusion tasks, especially in nighttime scenarios

URL PDF HTML 收藏
1703.08001 2026-06-04 cs.CV cs.NA math.NA 78%

Nonlinear Spectral Image Fusion

非线性频谱图像融合

Martin Benning, Michael Möller, Raz Z. Nossek, Martin Burger, Daniel Cremers, Guy Gilboa, Carola-Bibiane Schönlieb

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文展示基于总变分正则化的非线性频谱分解框架在图像融合及更广泛的图像处理任务中的有效性,通过选择特定图像的频率转移特征如面部皱纹,实现图像编辑。

Comments 13 pages, 9 figures, submitted to SSVM conference proceedings 2017

详情
AI中文摘要

本文演示了基于总变分正则化的非线性频谱分解框架在图像融合及更广泛的图像处理任务中的有效性。局部化良好且边缘保留的频谱总变分分解允许选择特定图像的频率以转移特定特征,如面部皱纹,从一个图像到另一个图像。我们通过多个数值实验展示了所提出方法的有效性,包括与泊松图像编辑、线性渗透、小波融合和拉普拉斯金字塔融合等竞争技术的比较。我们得出结论,所提出的频谱总变分图像分解框架是半自动和全自动图像编辑和融合的重要工具。

英文摘要

In this paper we demonstrate that the framework of nonlinear spectral decompositions based on total variation (TV) regularization is very well suited for image fusion as well as more general image manipulation tasks. The well-localized and edge-preserving spectral TV decomposition allows to select frequencies of a certain image to transfer particular features, such as wrinkles in a face, from one image to another. We illustrate the effectiveness of the proposed approach in several numerical experiments, including a comparison to the competing techniques of Poisson image editing, linear osmosis, wavelet fusion and Laplacian pyramid fusion. We conclude that the proposed spectral TV image decomposition framework is a valuable tool for semi- and fully-automatic image editing and fusion.

URL PDF HTML 收藏
2602.01760 2026-05-22 cs.CV 78%

MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement

MagicFuse: 单图像融合用于视觉与语义增强

Hao Zhang, Yanping Zha, Zizhuo Li, Meiqi Gong, Jiayi Ma

机构 * Electronic Information School, Wuhan University, China(武汉大学电子信息学院) Suzhou Institute of Wuhan University, China(武汉大学苏州研究院) School of Automation, Wuhan University, China(武汉大学自动化学院)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出MagicFuse单图像融合框架,通过扩散模型生成跨光谱场景表示,实现视觉与语义的双重约束,实验表明其性能优于多模态融合方法。

Comments Accepted by CVPR 2026

详情
AI中文摘要

本文聚焦于一个高度实用的场景:在仅使用可见成像传感器的情况下,如何继续利用多模态图像融合的优势。为此,我们提出了一种新的单图像融合概念,将其扩展到知识层面。具体而言,我们开发了MagicFuse,一种新的单图像融合框架,能够从单个低质量可见图像中推导出全面的跨光谱场景表示。MagicFuse首先引入了基于扩散模型的内在光谱知识增强分支和跨光谱知识生成分支。它们分别挖掘在可见光谱中被掩盖的场景信息,并学习转移到红外光谱的热辐射分布模式。在此基础上,我们设计了一个多领域知识融合分支,整合这两个分支的扩散流的概率噪声,从而通过连续采样获得跨光谱场景表示。然后,我们施加了视觉和语义约束,确保该场景表示能够满足人类观察同时支持下游语义决策。大量实验表明,尽管仅依赖单个退化的可见图像,我们的MagicFuse在视觉和语义表示性能上与或优于多模态输入的最先进融合方法。代码已公开在https://github.com/zhayanping/MagicFuse。

英文摘要

This paper focuses on a highly practical scenario: how to continue benefiting from the advantages of multi-modal image fusion under harsh conditions when only visible imaging sensors are available. To achieve this goal, we propose a novel concept of single-image fusion, which extends conventional data-level fusion to the knowledge level. Specifically, we develop MagicFuse, a novel single image fusion framework capable of deriving a comprehensive cross-spectral scene representation from a single low-quality visible image. MagicFuse first introduces an intra-spectral knowledge reinforcement branch and a cross-spectral knowledge generation branch based on the diffusion models. They mine scene information obscured in the visible spectrum and learn thermal radiation distribution patterns transferred to the infrared spectrum, respectively. Building on them, we design a multi-domain knowledge fusion branch that integrates the probabilistic noise from the diffusion streams of these two branches, from which a cross-spectral scene representation can be obtained through successive sampling. Then, we impose both visual and semantic constraints to ensure that this scene representation can satisfy human observation while supporting downstream semantic decision-making. Extensive experiments show that our MagicFuse achieves visual and semantic representation performance comparable to or even better than state-of-the-art fusion methods with multi-modal inputs, despite relying solely on a single degraded visible image. The code is publicly available at https://github.com/zhayanping/MagicFuse.

URL PDF HTML 收藏
2605.09455 2026-05-12 cs.CV 78%

Adaptive 3D Convolution for Remote Sensing Image Fusion

自适应3D卷积用于遥感图像融合

Siran Peng, Xiangyu Zhu, Shang-Qi Deng, Liang-Jian Deng, Zhen Lei

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Mathematical Sciences/Multi-Hazard Early Warning Key Laboratory of Sichuan Province, University of Electronic Science and Technology of China(数学科学学院/四川省多灾种早期预警重点实验室,电子科技大学) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(人工智能与机器人中心,香港科学院,中国科学院)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出自适应3D卷积方法,通过生成内容感知的3D卷积核,有效整合空间和光谱信息,提升遥感图像融合的精度与效率。

Comments Accepted by IEEE Transactions on Image Processing (TIP), Early Access, 2026

详情
AI中文摘要

遥感图像融合旨在从高分辨率图像(光谱信息有限)和低分辨率图像(光谱数据丰富)中生成高分辨率多/超光谱图像。近年来,深度学习技术在该领域表现出显著成效。大多数基于深度学习的方法将图像融合视为二维问题,通过将光谱信息编码到特征图通道中。然而,我们的研究发现这种策略引入了显著的光谱失真。相反,一些方法将光谱数据视为额外维度,利用标准3D卷积来保留光谱信息。然而,在标准3D卷积层中,相同的内核被应用于所有输入区域,这在图像融合中被发现是次优的。此外,标准3D卷积需要大量计算资源。为了解决这些挑战,我们提出了一种新的卷积范式,称为自适应3D卷积(Ada3D),用于遥感图像融合。Ada3D为每个输入体素应用独特的3D内核,能够捕捉细粒度细节。这些自适应内核通过两步过程生成:(i) 空间和光谱内核分别从各自图像源中导出;(ii) 这两种类型的内核然后结合形成内容感知的3D内核,有效整合空间和光谱信息。此外,引入了自适应偏置以增强卷积结果在体素层面。此外,我们还结合了组卷积技术以减少计算复杂性。因此,Ada3D以高效的方式实现了全面的自适应性。在五个数据集上的评估结果表明,我们的方法实现了SOTA性能,凸显了Ada3D的优越性。代码可在https://github.com/PSRben/Ada3D获取。

英文摘要

Remote sensing image fusion aims to create a high-resolution multi/hyper-spectral image from a high-resolution image with limited spectral information and a low-resolution image with abundant spectral data. Recently, deep learning (DL) techniques have shown significant effectiveness in this area. Most DL-based methods approach image fusion as a 2D problem by encoding spectral information into feature map channels. However, our research suggests that this strategy introduces notable spectral distortions. In contrast, some methods consider spectral data as an additional dimension, utilizing standard 3D convolutions to preserve spectral information. Nevertheless, in a standard 3D convolutional layer, the same set of kernels is applied across all input regions, which we have found to be sub-optimal for image fusion. Furthermore, standard 3D convolutions necessitate substantial computational resources. To address these challenges, we propose a novel convolutional paradigm called Adaptive 3D Convolution (Ada3D) for remote sensing image fusion. Ada3D applies a unique set of 3D kernels to each input voxel, enabling the capture of fine-grained details. These adaptive kernels are generated through a two-step process: (i) spatial and spectral kernels are derived from their respective image sources; (ii) these two types of kernels are then combined to form content-aware 3D kernels that effectively integrate spatial and spectral information. Additionally, adaptive biases are introduced to enhance the convolutional outcome at the voxel level. Furthermore, we incorporate the group convolution technique to reduce computational complexity. As a result, Ada3D offers full adaptivity in an efficient manner. Evaluation results across five datasets demonstrate that our method achieves SOTA performance, underscoring the superiority of Ada3D. The code is available at https://github.com/PSRben/Ada3D.

URL PDF HTML 收藏
2605.06969 2026-05-12 cs.CV 78%

Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment

将多模态大语言模型引入红外-可见图像融合质量评估

Yuchen Guo, Junli Gong, Yao Lu, Xintong Xu, Yiuming Cheung, Weifeng Su

机构 * Northwestern University(西北大学) Northeastern University(东北大学) University of Washington(华盛顿大学) Hong Kong Baptist University(香港 Baptist大学) Beijing Normal - Hong Kong Baptist University(北京师范大学-香港 Baptist大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出FuScore,利用多模态大语言模型生成连续质量评分,以更精确区分质量相近的融合图像,并结合多维子维度一致性和三重目标函数提升评估性能。

详情
AI中文摘要

红外-可见图像融合(IVIF)旨在将热信息和详细空间结构整合到单一融合图像中以增强感知。然而,现有评估方法倾向于过度优化手工制作的无参考统计和全参考度量,这些方法将源图像视为伪地面真实。最近的IVIF奖励建模努力从人类评分中学习,但使用聚合分数的标量回归,既未利用多模态大语言模型(MLLMs)的推理能力,也未在监督中编码每张图像的感知模糊性。为了应对这一问题,我们引入FuScore,利用MLLM生成连续质量评分,而非离散级别预测,从而在质量相近的融合图像之间实现细粒度区分。我们利用四个特定于IVIF的子维度的一致性来构建每张图像的软标签,其锐度反映整体判断的一致性。我们进一步引入一个三重目标,结合每张图像的分布监督、源对内Thurstone保真度用于方法级排序,以及跨源对Thurstone保真度用于场景级排序。大量实验表明,FuScore在与人类视觉偏好相关性方面达到了最先进的水平。

英文摘要

Infrared-Visible image fusion (IVIF) aims to integrate thermal information and detailed spatial structures into a single fused image to enhance perception. However, existing evaluation approaches tend to over-optimize both hand-crafted no-reference statistics and full-reference metrics that treat the source images as pseudo ground truths. Recent IVIF reward-modelling efforts learn from human ratings but use scalar regression on aggregated scores, neither leveraging the reasoning of Multimodal Large Language Models (MLLMs) nor encoding per-image perceptual ambiguity in their supervision, but naively introducing MLLMs with discrete one-hot supervision likewise collapses fused images of similar quality into different rating levels. To address this, we introduce FuScore, which utilizes an MLLM to mimic human visual perception by producing continuous quality score, rather than discrete level predictions, enabling fine-grained discrimination among fused images of similar quality. We exploit the agreement among four IVIF-specific sub-dimensions to construct a per-image soft label whose sharpness reflects how consensual the overall judgment is. We further introduce a tripartite objective combining per-image distributional supervision, within-source-pair Thurstone fidelity for method-level ordering, and cross-source-pair Thurstone fidelity for scene-level ordering across scenes. Extensive experiments demonstrate that FuScore achieves state-of-the-art correlation with human visual preferences.

URL PDF HTML 收藏
2605.06049 2026-05-08 cs.CV 78%

Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization

融合你的方式:通过直接偏好优化对齐图像融合与异质需求

Weijian Su, Songqian Zhang, Yuqi Han, Jian Zhuang, Yongdong Huang, Qiang Zhang

机构 * School of Computer Science and Technology, Dalian University of Technology(大连理工大学计算机科学与技术学院) Key Laboratory of Social Computing and Cognitive Intelligence (Dalian University of Technology), Ministry of Education(社会计算与认知智能重点实验室(大连理工大学),教育部) Institute of Image Processing and Understanding, North Minzu University(北华大学图像处理与理解研究所)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出DPOFusion框架,通过整合PALDM和PCLDM实现任务引导和偏好适应的图像融合,解决人类和机器视觉的异质需求问题,提升融合质量和任务导向的迁移能力。

Comments Accepted by CVPR 2026

详情
AI中文摘要

作为多模态处理的关键技术,红外与可见图像融合(IVIF)在整合互补光谱信息以增强视觉效果和下游视觉任务中起关键作用。尽管取得显著进展,现有方法难以灵活满足异质需求。实现适应性融合以符合人类和机器视觉的各种偏好仍是一个开放且具有挑战性的问题。为了解决这一挑战,我们提出DPOFusion,一种整合属性对齐潜在扩散模型(PALDM)和偏好可控潜在扩散模型(PCLDM)的直接偏好优化(DPO)框架,实现任务引导和偏好适应的IVIF,适用于人类和机器视觉。PALDM利用潜在融合先验和联合条件损失生成具有各种属性的多样化候选融合结果。PCLDM随后通过实例直接偏好优化(IDPO)进行微调,实现通过异质偏好信号直接控制最终融合结果。实验结果表明,我们的框架不仅实现了人类、视觉语言模型和任务驱动网络之间的精确偏好对齐,还为适应性融合质量和任务导向的迁移能力设定了新的基准。

英文摘要

As a key technique in multi-modal processing, infrared and visible image fusion (IVIF) plays a crucial role in integrating complementary spectral information for visual enhancement and downstream vision tasks. Despite remarkable progress, existing methods struggle to flexibly accommodate heterogeneous demands. Achieving adaptive fusion that aligns with various preferences from both human and machine vision remains an open and challenging problem. To address this challenge, we propose DPOFusion, a direct preference optimization (DPO) framework integrating the property-aligned latent diffusion model (PALDM) and the preference-controllable latent diffusion model (PCLDM), enabling task-guided, preference-adaptive IVIF for both human and machine vision. The PALDM leverages a latent fusion prior and a joint conditional loss to generate diverse candidate fusion results with various properties. PCLDM is subsequently fine-tuned via instance direct preference optimization (IDPO), enabling direct control of the final fusion results with heterogeneous preference signals. Experimental results demonstrate that our framework not only attains precise preference alignment among humans, vision-language models, and task-driven networks, but also sets a new benchmark for adaptive fusion quality and task-oriented transferability.

URL PDF HTML 收藏
2605.00885 2026-05-05 cs.CV 78%

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion

多分支非均匀图像去雾 via 浓度分区和图像融合

Yingming Zhang, Wuqi Su, Qing Xiao, Yonggang Yang

机构 * School of Software and Internet of Things Engineering, Zhejiang Gongshang University(软件与物联网工程学院,浙江工商大学) School of Computer Science and Technology, Tiangong University(计算机科学与技术学院,天工大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出CPIFNet网络,通过将非均匀去雾问题分解为多个均匀子问题,结合图像增强和融合技术,提升复杂场景下的去雾效果。

详情
AI中文摘要

现有单图像去雾方法在均匀薄雾图像上表现良好,但难以处理非均匀雾气浓度变化和区域密度突变的图像。为此,本文提出CPIFNet,通过将图像分解为多个局部区域,每个区域近似均匀雾气特征,采用两阶段架构:图像增强网络(IENet)和图像融合网络(IFNet)。IENet分支独立训练于不同浓度水平的均匀雾气数据集,生成擅长恢复对应雾气密度区域的增强模型。IFNet通过深度特征堆叠和融合,整合所有增强输出的优势区域,生成高质量去雾结果。此外,引入包含重建、感知、结构和颜色损失的综合损失函数,联合监督两个阶段。

英文摘要

Existing single image dehazing methods have demonstrated satisfactory performance on homogeneous thin-haze images; however, they often struggle with non-homogeneous hazy images that exhibit spatially varying haze concentrations and abrupt density transitions across different regions. To address this fundamental limitation, we propose a novel multi-branch deep neural network framework, termed Concentration Partitioning and Image Fusion Network (CPIFNet), which decomposes the challenging non-homogeneous dehazing problem into a set of tractable homogeneous sub-problems. Our key insight is that a single non-homogeneous hazy image can be viewed as a composite of multiple local regions, each exhibiting approximately homogeneous haze characteristics. CPIFNet employs a two-stage architecture consisting of an Image Enhancement Network (IENet) stage and an Image Fusion Network (IFNet) stage. In the first stage, multiple IENet branches are independently trained on homogeneous haze datasets of different concentration levels, producing enhancement models that excel at restoring regions matching their respective haze densities. In the second stage, the IFNet intelligently aggregates the advantageous regions from all enhancement outputs through deep feature stacking and merging, yielding a unified high-quality dehazed result. Furthermore, we introduce a comprehensive loss function incorporating reconstruction, perceptual, structural, and color losses to jointly supervise both stages.

URL PDF HTML 收藏
2604.10584 2026-04-15 cs.CV 78%

CoFusion: Multispectral and Hyperspectral Image Fusion via Spectral Coordinate Attention

CoFusion:通过光谱坐标注意力实现多谱段和超谱图像融合

Baisong Li

机构 * School of Integrated Circuits, Tsinghua University, Beijing 100084, China(集成电路学院,清华大学,北京100084,中国)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出CoFusion框架,通过建模跨尺度和跨模态依赖,提升多谱段和超谱图像融合的空谱协同效果,实现空谱细节增强与光谱保真度的平衡。

详情
AI中文摘要

多谱段和超谱图像融合(MHIF)旨在通过整合低分辨率超谱图像(LRHSI)和高分辨率多谱段图像(HRMSI)重建高分辨率图像。然而,现有方法在建模跨尺度交互和空谱协同方面存在局限,难以在空谱细节增强和光谱保真度之间取得最佳平衡。为此,我们提出CoFusion:一种统一的空谱协同融合框架,明确建模跨尺度和跨模态依赖。具体而言,设计了一个多尺度生成器(MSG),构建三级金字塔架构,有效整合全局语义和局部细节。在每个尺度内,采用双分支策略:空间坐标感知混合模块(SpaCAM)用于捕捉多尺度空间上下文,而光谱坐标感知混合模块(SpeCAM)通过频域分解和坐标混合增强光谱表示。此外,我们引入了空间-光谱交叉融合模块(SSCFM),用于动态跨模态对齐和互补特征融合。在多个基准数据集上的广泛实验表明,CoFusion在空间重建和光谱一致性方面均优于现有最先进方法。

英文摘要

Multispectral and Hyperspectral Image Fusion (MHIF) aims to reconstruct high-resolution images by integrating low-resolution hyperspectral images (LRHSI) and high-resolution multispectral images (HRMSI). However, existing methods face limitations in modeling cross-scale interactions and spatial-spectral collaboration, making it difficult to achieve an optimal trade-off between spatial detail enhancement and spectral fidelity. To address this challenge, we propose CoFusion: a unified spatial-spectral collaborative fusion framework that explicitly models cross-scale and cross-modal dependencies. Specifically, a Multi-Scale Generator (MSG) is designed to construct a three-level pyramidal architecture, enabling the effective integration of global semantics and local details. Within each scale, a dual-branch strategy is employed: the Spatial Coordinate-Aware Mixing module (SpaCAM) is utilized to capture multi-scale spatial contexts, while the Spectral Coordinate-Aware Mixing module (SpeCAM) enhances spectral representations through frequency decomposition and coordinate mixing. Furthermore, we introduce the Spatial-Spectral Cross-Fusion Module (SSCFM) to perform dynamic cross-modal alignment and complementary feature fusion. Extensive experiments on multiple benchmark datasets demonstrate that CoFusion consistently outperforms state-of-the-art methods, achieving superior performance in both spatial reconstruction and spectral consistency.

URL PDF HTML 收藏
2406.13621 2026-04-14 cs.CL cs.CV cs.LG 78%

LaMI: Augmenting Large Language Models via Late Multi-Image Fusion

LaMI:通过晚期多图像融合增强大型语言模型

Guy Yariv, Idan Schwartz, Yossi Adi, Sagie Benaim

机构 * The Hebrew University of Jerusalem(耶路撒冷希伯来大学) Bar-Ilan University(巴伊兰大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 LaMI通过在测试时融合多个图像提升视觉常识推理能力,同时保持文本推理性能,适用于视觉和NLP任务。

Comments Accepted to ACL 2026

详情
AI中文摘要

常识推理常需要文本和视觉知识,但仅训练于文本的大型语言模型(LLMs)缺乏视觉 grounding(例如,'帝企鹅腹部的颜色是什么?')。视觉语言模型(VLMs)在视觉任务上表现更好,但面临两个限制:(i)在文本常识推理上表现不如文本训练的LLMs;(ii)适应新发布的LLMs通常需要昂贵的多模态训练。另一种方法是在测试时向LLMs添加视觉信号,提升视觉常识推理而不损害文本推理能力,但先前设计通常依赖于早期融合和单张图像,可能效果不佳。我们提出了一种晚期多图像融合方法:通过轻量级并行采样生成多个图像,将这些图像的预测概率与文本-only LLM的预测概率通过一个晚期融合层结合,该层在最终预测前整合投影的视觉特征。在视觉常识和NLP基准测试中,我们的方法在视觉推理上显著优于增强的LLMs,在基于视觉的任务上与VLMs相当,并且当应用于强大的LLMs如LLaMA 3时,也能提升NLP性能,同时仅增加少量的测试时间开销。项目页面可在:https://guyyariv.github.io/LaMI访问。

英文摘要

Commonsense reasoning often requires both textual and visual knowledge, yet Large Language Models (LLMs) trained solely on text lack visual grounding (e.g., "what color is an emperor penguin's belly?"). Visual Language Models (VLMs) perform better on visually grounded tasks but face two limitations: (i) often reduced performance on text-only commonsense reasoning compared to text-trained LLMs, and (ii) adapting newly released LLMs to vision input typically requires costly multimodal training. An alternative augments LLMs with test-time visual signals, improving visual commonsense without harming textual reasoning, but prior designs often rely on early fusion and a single image, which can be suboptimal. We propose a late multi-image fusion method: multiple images are generated from the text prompt with a lightweight parallel sampling, and their prediction probabilities are combined with those of a text-only LLM through a late-fusion layer that integrates projected visual features just before the final prediction. Across visual commonsense and NLP benchmarks, our method significantly outperforms augmented LLMs on visual reasoning, matches VLMs on vision-based tasks, and, when applied to strong LLMs such as LLaMA 3, also improves NLP performance while adding only modest test-time overhead. Project page is available at: https://guyyariv.github.io/LaMI.

URL PDF HTML 收藏
2604.09030 2026-04-13 cs.CV 78%

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Multi-Exposure Image Fusion in Dynamic Scenes (Track 2)

NTIRE 2026 第三次恢复任何图像模型(RAIM)挑战:动态场景中的多曝光图像融合(Track 2)

Lishen Qu, Yao Liu, Jie Liang, Hui Zeng, Wen Dai, Guanyi Qin, Ya-nan Guan, Shihao Zhou, Jufeng Yang, Lei Zhang, Radu Timofte, Xiyuan Yuan, Wanjie Sun, Shihang Li, Bo Zhang, Bin Chen, Jiannan Lin, Yuxu Chen, Qinquan Gao, Tong Tong, Song Gao, Jiacong Tang, Tao Hu, Xiaowen Ma, Qingsen Yan, Sunhan Xu, Juan Wang, Xinyu Sun, Lei Qi, He Xu, Jiachen Tu, Guoyi Xu, Yaoxin Jiang, Jiajia Liu, Yaokun Shi

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出NTIRE 2026挑战,针对动态场景中多曝光图像融合的HDR成像难题,提供包含100训练序列和100测试序列的基准数据集,评估PSNR、SSIM、LPIPS等指标及感知质量,最终由114支队伍提交987份方案,提升去噪和细节恢复能力。

Comments Accepted by CVPRW 2026

详情
AI中文摘要

本文介绍了NTIRE 2026,即第三次恢复任何图像模型(RAIM)挑战,聚焦于动态场景中的多曝光图像融合。我们引入了一个针对实际但困难的HDR成像设置的基准,其中必须在场景运动、光照变化和手持相机抖动的情况下进行曝光拼接。挑战数据集包含100个训练序列(7个曝光级别)和100个测试序列(5个曝光级别),反映了现实场景中常导致对齐错误和鬼影伪影的情况。我们通过PSNR、SSIM和LPIPS导出的排行榜分数评估提交方案,同时考虑感知质量、效率和可重复性。该赛道吸引了114支参赛队伍,共收到987份提交。获胜方法显著提升了从多曝光融合中去除伪影和恢复细节的能力。数据集和各团队代码可在仓库:https://github.com/qulishen/RAIM-HDR 中找到。

英文摘要

This paper presents NTIRE 2026, the 3rd Restore Any Image Model (RAIM) challenge on multi-exposure image fusion in dynamic scenes. We introduce a benchmark that targets a practical yet difficult HDR imaging setting, where exposure bracketing must be fused under scene motion, illumination variation, and handheld camera jitter. The challenge data contains 100 training sequences with 7 exposure levels and 100 test sequences with 5 exposure levels, reflecting real-world scenarios that frequently cause misalignment and ghosting artefacts. We evaluate submissions with a leaderboard score derived from PSNR, SSIM, and LPIPS, while also considering perceptual quality, efficiency, and reproducibility during the final review. This track attracted 114 participating teams and received 987 submissions. The winning methods significantly improved the ability to remove artifacts from multi-exposure fusion and recover fine details. The dataset and the code of each team can be found at the repository: https://github.com/qulishen/RAIM-HDR.

URL PDF HTML 收藏
2604.08924 2026-04-13 cs.CV 78%

Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image Fusion

定制融合:一种闭环动态网络用于自适应多任务感知红外-可见图像融合

Zengyi Yang, Yu Liu, Juan Cheng, Zhiqin Zhu, Yafei Zhang, Huafeng Li

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出闭环动态网络CLDyN,通过需求驱动语义补偿模块实现多任务自适应的图像融合,实验表明其在保持高质量融合的同时具备强多任务适应性。

Comments This paper has been accepted by CVPR 2026

详情
AI中文摘要

红外-可见图像融合旨在整合互补信息以实现稳健的视觉理解,但现有融合方法难以同时适应多种下游任务。为了解决这一问题,我们提出了一种闭环动态网络(CLDyN),能够根据多样下游任务的语义需求进行自适应的图像融合。具体而言,CLDyN引入了一个闭环优化机制,通过需求驱动语义补偿(RSC)模块建立语义传输链,实现从下游任务到融合网络的显式反馈。RSC模块利用基向量库(BVB)和架构适应性语义注入(A2SI)块,根据任务需求定制网络架构,从而实现任务特定的语义补偿,使融合网络能够主动适应多种任务而无需重新训练。为了促进语义补偿,引入了奖励-惩罚策略,根据任务性能变化奖励或惩罚RSC模块。在M3FD、FMB和VT5000数据集上的实验表明,CLDyN不仅保持了高质量的融合效果,还表现出强大的多任务适应能力。代码可在https://github.com/YR0211/CLDyN获取。

英文摘要

Infrared-visible image fusion aims to integrate complementary information for robust visual understanding, but existing fusion methods struggle with simultaneously adapting to multiple downstream tasks. To address this issue, we propose a Closed-Loop Dynamic Network (CLDyN) that can adaptively respond to the semantic requirements of diverse downstream tasks for task-customized image fusion. Specifically, CLDyN introduces a closed-loop optimization mechanism that establishes a semantic transmission chain to achieve explicit feedback from downstream tasks to the fusion network through a Requirement-driven Semantic Compensation (RSC) module. The RSC module leverages a Basis Vector Bank (BVB) and an Architecture-Adaptive Semantic Injection (A2SI) block to customize the network architecture according to task requirements, thereby enabling task-specific semantic compensation and allowing the fusion network to actively adapt to diverse tasks without retraining. To promote semantic compensation, a reward-penalty strategy is introduced to reward or penalize the RSC module based on task performance variations. Experiments on the M3FD, FMB, and VT5000 datasets demonstrate that CLDyN not only maintains high fusion quality but also exhibits strong multi-task adaptability. The code is available at https://github.com/YR0211/CLDyN.

URL PDF HTML 收藏
2604.08922 2026-04-13 cs.CV 78%

Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios

抗退化融合:一种高效的抗退化扩散框架,用于任意退化场景下的多模态图像融合

Yu Shi, Yu Liu, Zhong-Cheng Wu, Juan Cheng, Huafeng Li, Xun Chen

机构 * Department of Biomedical Engineering, Hefei University of Technology(合肥工业大学生物医学工程系) Faculty of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学技术学院)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出一种高效的抗退化扩散框架,用于处理多模态图像融合中的复杂退化场景。该框架通过隐式去噪和联合观测模型校正机制,提升了在复杂退化下的融合性能。

Comments Accepted by CVPR 2026

详情
AI中文摘要

复杂的退化如噪声、模糊和低分辨率是现实世界图像融合任务中的典型挑战,限制了现有方法的性能和实用性。端到端神经网络方法通常易于设计且推理高效,但其黑盒性质导致可解释性有限。基于扩散的方法在一定程度上缓解了这一问题,通过提供强大的生成先验和更结构化的推理过程。然而,它们被训练以学习单一域的目标分布,而融合缺乏自然融合数据,依赖于从多个来源建模互补信息,使扩散在实践中难以直接应用。为了解决这些挑战,本文提出了一种高效的抗退化扩散框架,用于在任意退化场景下的图像融合。具体而言,与传统扩散模型显式预测噪声不同,我们的方法通过直接回归融合图像进行隐式去噪,使在复杂退化下灵活适应多种融合任务成为可能,且仅需有限步骤。此外,我们设计了一个联合观测模型校正机制,在采样过程中同时施加退化和融合约束,以确保高重建精度。在多样化的融合任务和退化配置上的实验表明,所提出的方法在复杂退化场景下具有优越性。

英文摘要

Complex degradations like noise, blur, and low resolution are typical challenges in real world image fusion tasks, limiting the performance and practicality of existing methods. End to end neural network based approaches are generally simple to design and highly efficient in inference, but their black-box nature leads to limited interpretability. Diffusion based methods alleviate this to some extent by providing powerful generative priors and a more structured inference process. However, they are trained to learn a single domain target distribution, whereas fusion lacks natural fused data and relies on modeling complementary information from multiple sources, making diffusion hard to apply directly in practice. To address these challenges, this paper proposes an efficient degradation aware diffusion framework for image fusion under arbitrary degradation scenarios. Specifically, instead of explicitly predicting noise as in conventional diffusion models, our method performs implicit denoising by directly regressing the fused image, enabling flexible adaptation to diverse fusion tasks under complex degradations with limited steps. Moreover, we design a joint observation model correction mechanism that simultaneously imposes degradation and fusion constraints during sampling to ensure high reconstruction accuracy. Experiments on diverse fusion tasks and degradation configurations demonstrate the superiority of the proposed method under complex degradation scenarios.

URL PDF HTML 收藏
2604.05742 2026-04-08 cs.CV 78%

ASSR-Net: Anisotropic Structure-Aware and Spectrally Recalibrated Network for Hyperspectral Image Fusion

ASSR-Net:一种用于超分辨率图像融合的各向异性结构感知与频谱重校准网络

Qiya Song, Hongzhi Zhou, Lishan Tan, Renwei Dian, Shutao Li

机构 * School of Information Science and Engineering, Hunan Normal University(湖南师范大学信息科学与工程学院) School of Robotics, Hunan University(湖南大学机器人学院)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 ASSR-Net通过两阶段融合策略,解决超分辨率图像融合中各向异性结构重建不足和频谱失真问题,提升空间细节和频谱一致性。

详情
AI中文摘要

超分辨率图像融合旨在通过多源输入整合互补信息来重建高空间分辨率的超分辨率图像(HR-HSI)。尽管近年来取得了进展,现有方法仍面临两个关键挑战:(1)各向异性空间结构重建不足,导致细节模糊和空间质量下降;(2)融合过程中的频谱失真,阻碍精细频谱表示。为此,我们提出了ASSR-Net:一种用于超分辨率图像融合的各向异性结构感知与频谱重校准网络。ASSR-Net采用两阶段融合策略,包括各向异性结构感知空间增强(ASSE)和分层先验引导频谱校准(HPSC)。第一阶段中,方向感知融合模块自适应地捕获多方向的结构特征,有效重建各向异性空间模式。第二阶段中,频谱重校准模块利用原始低分辨率HSI作为频谱先验,显式纠正融合结果中的频谱偏差,从而提升频谱保真度。在各种基准数据集上的广泛实验表明,ASSR-Net在各种基准数据集上均优于现有最先进方法,实现了更优的空间细节保留和频谱一致性。

英文摘要

Hyperspectral image fusion aims to reconstruct high-spatial-resolution hyperspectral images (HR-HSI) by integrating complementary information from multi-source inputs. Despite recent progress, existing methods still face two critical challenges: (1) inadequate reconstruction of anisotropic spatial structures, resulting in blurred details and compromised spatial quality; and (2) spectral distortion during fusion, which hinders fine-grained spectral representation. To address these issues, we propose \textbf{ASSR-Net}: an Anisotropic Structure-Aware and Spectrally Recalibrated Network for Hyperspectral Image Fusion. ASSR-Net adopts a two-stage fusion strategy comprising anisotropic structure-aware spatial enhancement (ASSE) and hierarchical prior-guided spectral calibration (HPSC). In the first stage, a directional perception fusion module adaptively captures structural features along multiple orientations, effectively reconstructing anisotropic spatial patterns. In the second stage, a spectral recalibration module leverages the original low-resolution HSI as a spectral prior to explicitly correct spectral deviations in the fused results, thereby enhancing spectral fidelity. Extensive experiments on various benchmark datasets demonstrate that ASSR-Net consistently outperforms state-of-the-art methods, achieving superior spatial detail preservation and spectral consistency.

URL PDF HTML 收藏
2604.02896 2026-04-06 cs.CV 78%

EvaNet: Towards More Efficient and Consistent Infrared and Visible Image Fusion Assessment

EvaNet:迈向更高效且一致的红外与可见图像融合评估

Chunyang Cheng, Tianyang Xu, Xiao-Jun Wu, Tao Zhou, Hui Li, Zhangyong Tang, Josef Kittler

机构 * Jiangnan University(江南大学) University of Surrey(萨里大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出EvaNet框架,通过轻量网络高效近似常用指标,结合对比学习和大语言模型感知评估,实现更一致的图像融合评估。

Comments 20 figures,accepted by TPAMI

详情
AI中文摘要

在图像融合研究中,评估至关重要,但现有指标大多直接借鉴其他视觉任务而未适配。这些传统指标基于复杂图像变换,不仅无法捕捉融合结果的真实质量,还计算成本高。为此,我们提出一个专门针对图像融合的统一评估框架。其核心是一个轻量网络,通过分而治之策略高效近似常用指标。不同于传统方法直接评估融合图像与源图像的相似性,我们首先将融合结果分解为红外和可见成分,再通过评估模型测量这些分离成分的信息保留程度,从而有效解耦融合评估过程。训练时,我们结合对比学习策略,并利用大语言模型提供的感知场景评估信息来训练评估模型。最后,我们提出首个一致性评估框架,通过独立无参考分数和下游任务表现作为客观参考,测量图像融合指标与人类视觉感知的一致性。大量实验表明,我们的学习评估范式在多种标准图像融合基准上实现了更高的效率(高达1000倍)和一致性。我们的代码将在https://github.com/AWCXV/EvaNet上公开。

英文摘要

Evaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to capture the true quality of the fusion results but also are computationally demanding. To address these issues, we propose a unified evaluation framework specifically tailored for image fusion. At its core is a lightweight network designed efficiently to approximate widely used metrics, following a divide-and-conquer strategy. Unlike conventional approaches that directly assess similarity between fused and source images, we first decompose the fusion result into infrared and visible components. The evaluation model is then used to measure the degree of information preservation in these separated components, effectively disentangling the fusion evaluation process. During training, we incorporate a contrastive learning strategy and inform our evaluation model by perceptual scene assessment provided by a large language model. Last, we propose the first consistency evaluation framework, which measures the alignment between image fusion metrics and human visual perception, using both independent no-reference scores and downstream tasks performance as objective references. Extensive experiments show that our learning-based evaluation paradigm delivers both superior efficiency (up to 1,000 times faster) and greater consistency across a range of standard image fusion benchmarks. Our code will be publicly available at https://github.com/AWCXV/EvaNet.

URL PDF HTML 收藏
2604.01579 2026-04-03 cs.CV cs.AI 78%

Harmonized Tabular-Image Fusion via Gradient-Aligned Alternating Learning

通过梯度对齐交替学习实现表-图像融合

Longfei Huang, Yang Yang

机构 * Nanjing University of Science and Technology(南京理工大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出GAAL方法,通过对齐模态梯度解决多模态表-图像融合中的梯度冲突问题,提升融合性能。

Comments ICME 26

详情
AI中文摘要

多模态表-图像融合是一项新兴任务,近年来在多个领域受到越来越多的关注。然而,现有方法可能受到模态间梯度冲突的阻碍,误导单模态学习器的优化。本文提出了一种新的梯度对齐交替学习(GAAL)范式,通过对齐模态梯度来解决这一问题。具体而言,GAAL采用交替单模态学习和共享分类器,以解耦多模态梯度并促进交互。此外,我们设计了基于不确定性的跨模态梯度手术,以选择性地对齐跨模态梯度,从而引导共享参数以造福所有模态。结果表明,GAAL能够提供有效的单模态辅助,并帮助提升整体融合性能。通过在广泛使用的数据集上的实证实验,与各种最先进的(SoTA)表-图像融合基线和测试时表缺失基线进行比较,验证了我们方法的优越性。源代码可在https://github.com/njustkmg/ICME26-GAAL上获取。

英文摘要

Multimodal tabular-image fusion is an emerging task that has received increasing attention in various domains. However, existing methods may be hindered by gradient conflicts between modalities, misleading the optimization of the unimodal learner. In this paper, we propose a novel Gradient-Aligned Alternating Learning (GAAL) paradigm to address this issue by aligning modality gradients. Specifically, GAAL adopts an alternating unimodal learning and shared classifier to decouple the multimodal gradient and facilitate interaction. Furthermore, we design uncertainty-based cross-modal gradient surgery to selectively align cross-modal gradients, thereby steering the shared parameters to benefit all modalities. As a result, GAAL can provide effective unimodal assistance and help boost the overall fusion performance. Empirical experiments on widely used datasets reveal the superiority of our method through comparison with various state-of-the-art (SoTA) tabular-image fusion baselines and test-time tabular missing baselines. The source code is available at https://github.com/njustkmg/ICME26-GAAL.

URL PDF HTML 收藏
2510.24379 2026-04-03 cs.CV 78%

A Luminance-Aware Multi-Scale Network for Polarization Image Fusion with a Multi-Scene Dataset

一种考虑亮度的多尺度网络用于极化图像融合与多场景数据集

Zhuangfan Huang, Xiaosong Li, Gao Wang, Tao Ye, Haishu Tan, Huafeng Li

机构 * Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology, School of Physics and Optoelectronic Engineering, Foshan University(广东-香港-澳门智能微纳光电子技术联合实验室,物理与光电工程学院,佛山大学) State Key Laboratory of Dynamic Measurement Technology, North University of China(动态测试技术国家重点实验室,中北大学) School of Information Engineering and Automation, Kunming University of Science and Technology(信息工程与自动化学院,昆明理工大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出考虑亮度的多尺度网络,通过多尺度空间权重矩阵和全局-局部特征融合机制,提升极化图像融合在复杂光照环境下的性能,使用多场景数据集验证了方法的优越性。

详情
AI中文摘要

极化图像融合结合S0和DOLP图像,通过互补的纹理特征揭示表面粗糙度和材料属性,具有在伪装识别、组织病理学分析、表面缺陷检测等领域的应用价值。为整合不同极化图像在复杂亮度环境中的互补信息,我们提出一种亮度感知的多尺度网络(MLSN)。在编码阶段,我们通过亮度分支提出多尺度空间权重矩阵,动态加权将亮度注入特征图中,解决极化图像固有的对比度差异问题。在瓶颈层设计全局-局部特征融合机制,通过窗口自注意力计算,在特征维度重构阶段通过残差链接平衡全局上下文和局部细节。在解码阶段,为进一步提高对复杂光照的适应性,我们提出亮度增强模块,建立亮度分布与纹理特征之间的映射关系,实现融合结果的非线性亮度校正。我们还提出了MSP,一个包含1000对极化图像的多场景数据集,覆盖17种室内外复杂光照场景。MSP提供四个方向的极化原始地图,解决了极化图像融合中高质量数据集稀缺的问题。在MSP、PIF和GAND数据集上的大量实验验证,所提的MLSN在主观和客观评估中优于现有最先进方法,MS-SSIM和SD指标分别比其他方法的平均值高8.57%、60.64%、10.26%、63.53%、22.21%和54.31%。源代码和数据集可在https://github.com/1hzf/MLS-UNet获取。

英文摘要

Polarization image fusion combines S0 and DOLP images to reveal surface roughness and material properties through complementary texture features, which has important applications in camouflage recognition, tissue pathology analysis, surface defect detection and other fields. To intergrate coL-Splementary information from different polarized images in complex luminance environment, we propose a luminance-aware multi-scale network (MLSN). In the encoder stage, we propose a multi-scale spatial weight matrix through a brightness-branch , which dynamically weighted inject the luminance into the feature maps, solving the problem of inherent contrast difference in polarized images. The global-local feature fusion mechanism is designed at the bottleneck layer to perform windowed self-attention computation, to balance the global context and local details through residual linking in the feature dimension restructuring stage. In the decoder stage, to further improve the adaptability to complex lighting, we propose a Brightness-Enhancement module, establishing the mapping relationship between luminance distribution and texture features, realizing the nonlinear luminance correction of the fusion result. We also present MSP, an 1000 pairs of polarized images that covers 17 types of indoor and outdoor complex lighting scenes. MSP provides four-direction polarization raw maps, solving the scarcity of high-quality datasets in polarization image fusion. Extensive experiment on MSP, PIF and GAND datasets verify that the proposed MLSN outperms the state-of-the-art methods in subjective and objective evaluations, and the MS-SSIM and SD metircs are higher than the average values of other methods by 8.57%, 60.64%, 10.26%, 63.53%, 22.21%, and 54.31%, respectively. The source code and dataset is avalable at https://github.com/1hzf/MLS-UNet.

URL PDF HTML 收藏
2603.08018 2026-04-02 cs.CV 78%

Missing No More: Dictionary-Guided Cross-Modal Image Fusion under Missing Infrared

不再缺失:字典引导的跨模态图像融合在红外缺失情况下的应用

Yafei Zhang, Meng Ma, Huafeng Li, Yu Liu

机构 * Faculty of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院) Department of Biomedical Engineering, Hefei University of Technology(合肥工业大学生物医学工程系)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出了一种基于共享卷积字典的字典引导框架,解决红外缺失时的跨模态图像融合问题,通过联合字典学习、视觉引导红外推断和自适应融合方法提升感知质量和下游检测性能。

Comments This paper has been accepted by CVPR 2026

详情
AI中文摘要

红外-可见(IR-VIS)图像融合对于感知和安全至关重要,但大多数方法依赖于训练和推理时两种模态的可用性。当红外模态缺失时,像素空间生成替代品难以控制且缺乏可解释性。我们通过提出一个基于共享卷积字典的字典引导、系数域框架来解决缺失红外融合问题。该流程包括三个关键组件:(1)联合共享字典表示学习(JSRL)学习一个统一且可解释的原子空间,由IR和VIS模态共享;(2)视觉引导红外推断(VGII)将VIS系数转换为伪红外系数并在系数域中进行一次闭环细化,利用冻结的大语言模型作为弱语义先验;(3)通过窗口注意力和卷积混合在原子层面融合VIS结构和推断的红外线索,随后使用共享字典进行重建。该编码-转移-融合-重建流程避免了不可控的像素空间生成,同时在可解释的字典-系数表示中保持了先验信息。在缺失红外设置下的实验表明,感知质量和下游检测性能均有显著提升。据我们所知,这是首个联合学习共享字典并进行系数域推断-融合以解决缺失红外融合的框架。源代码可在https://github.com/harukiv/DCMIF公开获取。

英文摘要

Infrared-visible (IR-VIS) image fusion is vital for perception and security, yet most methods rely on the availability of both modalities during training and inference. When the infrared modality is absent, pixel-space generative substitutes become hard to control and inherently lack interpretability. We address missing-IR fusion by proposing a dictionary-guided, coefficient-domain framework built upon a shared convolutional dictionary. The pipeline comprises three key components: (1) Joint Shared-dictionary Representation Learning (JSRL) learns a unified and interpretable atom space shared by both IR and VIS modalities; (2) VIS-Guided IR Inference (VGII) transfers VIS coefficients to pseudo-IR coefficients in the coefficient domain and performs a one-step closed-loop refinement guided by a frozen large language model as a weak semantic prior; and (3) Adaptive Fusion via Representation Inference (AFRI) merges VIS structures and inferred IR cues at the atom level through window attention and convolutional mixing, followed by reconstruction with the shared dictionary. This encode-transfer-fuse-reconstruct pipeline avoids uncontrolled pixel-space generation while ensuring prior preservation within interpretable dictionary-coefficient representation. Experiments under missing-IR settings demonstrate consistent improvements in perceptual quality and downstream detection performance. To our knowledge, this represents the first framework that jointly learns a shared dictionary and performs coefficient-domain inference-fusion to tackle missing-IR fusion. The source code is publicly available at https://github.com/harukiv/DCMIF.

URL PDF HTML 收藏
2603.24296 2026-03-26 cs.CV 78%

AMIF: Authorizable Medical Image Fusion Model with Built-in Authentication

AMIF:具有内置认证的可授权医学图像融合模型

Jie Song, Jun Jia, Wei Sun, Wangqiu Zhou, Tao Tan, Guangtao Zhai

机构 * Macao Polytechnic University(澳门理工学院) Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学) Hefei University of Technology(合肥工业大学)

专题命中 Image Fusion :image fusion(title,abstract)

AI总结 本文提出AMIF,首个具有内置认证的医学图像融合模型,通过授权访问控制保护知识产权,防止推理泄露。

详情
AI中文摘要

多模态图像融合能够实现精确的病变定位和特征描述,从而提高诊断准确性,增强临床决策并推动其在医学影像研究中的重要性。强大的多模态图像融合模型依赖于高质量、临床代表性的多模态训练数据和精心设计的模型架构。因此,此类专业放射组学模型的开发是标准化采集、临床专业知识和算法设计能力的协作成果,需要保护相关知识产权。然而,当前多模态图像融合模型生成融合输出,但缺乏内置机制来保护知识产权,导致通过推理泄露暴露了专有模型知识和敏感训练数据。例如,恶意用户可以利用融合输出和模型蒸馏或其他基于推理的逆向工程技术来近似专有模型的融合性能。为了解决这个问题,我们提出了AMIF,首个具有内置认证的医学图像融合模型,将授权访问控制整合到图像融合目标中。对于未经授权的使用,AMIF将显式可见的版权标识嵌入融合结果中。相反,通过成功的关键基于认证可以访问高质量的融合结果。

英文摘要

Multimodal image fusion enables precise lesion localization and characterization for accurate diagnosis, thereby strengthening clinical decision-making and driving its growing prominence in medical imaging research. A powerful multimodal image fusion model relies on high-quality, clinically representative multimodal training data and a rigorously engineered model architecture. Therefore, the development of such professional radiomics models represents a collaborative achievement grounded in standardized acquisition, clinical-specific expertise, and algorithmic design proficiency, which necessitates protection of associated intellectual property rights. However, current multimodal image fusion models generate fused outputs without built-in mechanisms to safeguard intellectual property rights, inadvertently exposing proprietary model knowledge and sensitive training data through inference leakage. For example, malicious users can exploit fusion outputs and model distillation or other inference-based reverse engineering techniques to approximate the fusion performance of proprietary models. To address this issue, we propose AMIF, the first Authorizable Medical Image Fusion model with built-in authentication, which integrates authorization access control into the image fusion objective. For unauthorized usage, AMIF embeds explicit and visible copyright identifiers into fusion results. In contrast, high-quality fusion results are accessible upon successful key-based authentication.

URL PDF HTML 收藏