发表机构
Northwestern Polytechnical University; Hefei University of Technology; Rocket Force University of Engineering; Chongqing University of Posts and Telecommunications(西北工业大学; 合肥工业大学; 火箭军工程大学; 重庆邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对红外-可见光图像融合的信息差异挑战,提出基于双固有提示的P2Fusion框架,结合Teach-to-Fuse机制与GDER模块,在多数据集上实现SOTA性能并提升下游感知鲁棒性。
AI 中文摘要
红外-可见光图像融合(IVIF)对多模态感知至关重要,但协调热特征与纹理特征之间固有的信息差异仍是一项基本挑战。现有的先验引导方法常依赖静态约束,会引发优化冲突,或利用来自大规模基础模型(如CLIP、DINO)的外部语义先验,这类方法往往无法挖掘高保真融合所需的固有模态特性。为解决这些问题,我们提出P2Fusion,这是一种基于先验蒸馏的框架,通过双固有提示重新构建IVIF。我们没有施加硬编码惩罚,而是将图像固有先验(热显著性和空间质量)蒸馏为可学习的动态调节器。具体而言,一种“用于融合的教学”机制提供双粒度渐进式引导,结合门控动态专家重校准(GDER)模块实现解耦特征细化。该设计使网络能通过专家专业化自适应调节模态竞争。大量实验表明,P2Fusion在5个主流数据集上达到了最先进性能。值得注意的是,我们的框架在融合质量上展现出持续的性能优势,在5个基准的20项关键评估指标中,有14项取得了最先进结果。此外,它还有效提升了下游感知的鲁棒性,例如在MSRS数据集上目标检测的平均精度均值(mAP)提升了3.2%,在M3FD数据集上提升了0.5%,在DroneVehicle数据集上提升了0.9%。我们的代码将在该https网址发布。
英文摘要
Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.g., CLIP/DINO), which frequently fail to exploit the intrinsic modality characteristics essential for high-fidelity fusion. To address these issues, we propose P2Fusion, a prior-guided distillation-based framework that reformulates IVIF via dual intrinsic prompts. Instead of imposing hard-coded penalties, we distill image-intrinsic priors, thermal saliency and spatial quality, into learnable dynamic regulators. Specifically, a Teach-to-Fuse mechanism provides dual-granularity progressive guidance, coupled with a Gated Dynamic Expert Recalibration (GDER) module for decoupled feature refinement. This design enables the network to adaptively mediate modal competition through expert specialization. Extensive experiments demonstrate that P2Fusion achieves state-of-the-art performance across five mainstream datasets. Notably, our framework demonstrates consistent performance advantages in fusion quality, achieving state-of-the-art results in 14 out of 20 key evaluation metrics across 5 benchmarks. Furthermore, it effectively contributes to the robustness of downstream perception, such as +3.2% mAP on MSRS, +0.5% mAP on M3FD and +0.9% mAP on DroneVehicle for object detection. Our code will be available at https://github.com/YiShi99/P2Fusion
CommentsAccepted by ECCV 2026. Website: https://p2fusion.github.io