发表机构
College of Computing and Data Science, Nanyang Technological University; The Hong Kong Polytechnic University; Harbin Institute of Technology; State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(南洋理工大学计算与数据科学学院; 香港理工大学; 哈尔滨工业大学; 北京大学计算机科学学院多媒体信息处理国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究红外与可见光图像融合问题,提出CIS - Fuse脉冲网络,通过电流注入脉冲算子在膜电位水平实现跨模态融合,构建双向跨模态融合模块并部署在双分支架构上,实验表明其融合质量与基于ANN的方法相当且更节能。
AI 中文摘要
红外与可见光图像融合(IVIF)将两种模态的互补信息整合到具有更丰富场景内容的单个图像中。现有方法大多基于人工神经网络(ANNs),而脉冲神经网络(SNNs)通过稀疏二进制脉冲进行通信,仅在脉冲发生的时间和位置进行计算,为更节能的融合提供了途径。但直接将SNNs应用于IVIF存在根本矛盾:跨模态融合依赖于两种模态的细粒度响应,而二进制脉冲可能会丢弃低于激发阈值的互补线索。膜电位在激发前保留这些亚阈值响应,使两种模态在该阶段整合时共同塑造输出。基于此,我们提出了CIS - Fuse,一种直接在膜电位水平执行跨模态融合的脉冲网络。其核心是电流注入脉冲(CIS)算子,将一种模态作为门控辅助电流注入到另一种模态的驱动神经元中,使两者在脉冲激发前整合,具有逐通道可学习的注入强度来自适应调节调制幅度。基于CIS,我们构建了双向跨模态融合(BCMF)模块并将其部署在具有不对称堆叠深度的双分支架构上,两个分支具有明确的功能专业化。在四个IVIF基准测试以及下游检测和分割上的大量实验表明,CIS - Fuse在继承基于脉冲计算的能量效率的同时,实现了与基于ANN的最先进方法相当的融合质量,推理能量比类似规模的基于ANN的DCEvo低约一个数量级。代码将在发表后发布。
英文摘要
Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a single image with richer scene content. While existing methods are largely built on artificial neural networks (ANNs), which densely compute over all activations, spiking neural networks (SNNs) communicate through sparse binary spikes and compute only where and when a spike occurs, offering a route to more energy-efficient fusion. However, directly applying SNNs to IVIF creates a fundamental tension: cross-modal fusion relies on fine-grained responses from both modalities, whereas binary spikes can discard complementary cues that remain below the firing threshold. The membrane potential retains these subthreshold responses before firing, letting both modalities jointly shape the output when integrated at this stage. Building on this, we propose CIS-Fuse, a spiking network that performs cross-modal fusion directly at the membrane-potential level. At its core is the current injection spiking (CIS) operator, which injects one modality as a gated auxiliary current into the driving neuron of the other, so the two integrate before spike firing, with a per-channel learnable injection strength that adaptively regulates the modulation magnitude. Building on CIS, we construct a bidirectional cross-modal fusion (BCMF) module and deploy it on a dual-branch architecture with asymmetric stacking depths, where the two branches develop a clear functional specialization. Extensive experiments on four IVIF benchmarks and on downstream detection and segmentation show that CIS-Fuse achieves fusion quality on par with state-of-the-art ANN-based methods while inheriting the energy efficiency of spike-based computation, with roughly an order of magnitude lower inference energy than the similarly-sized ANN-based DCEvo. Code will be released upon publication.