发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对脉冲神经网络在图像-文本检索中跨模态语义结构捕捉不足的问题,提出MSGAT与Sim-Fuse策略,在Flickr30K、MSCOCO数据集上实现优于现有方法的性能,能耗降低55%。
AI 中文摘要
脉冲神经网络(SNN)通过稀疏的事件驱动计算提供高能效的计算范式,在高效多模态学习中展现出巨大潜力。然而,将SNN应用于图像-文本检索(ITR)这类高级多模态任务仍具挑战性,因为稀疏的脉冲表示难以捕捉跨模态对齐所需的语义结构。现有的脉冲ITR方法依赖局部对齐和训练期间的额外软标签监督,缺乏对结构和多粒度关系的感知。为解决这些问题,我们提出用于结构建模的多头脉冲图注意力网络(MSGAT),并为其配备动态注意力头以捕捉互补的关系模式,实现脉冲驱动的图推理与聚合。在双分支多粒度融合框架内,MSGAT生成的细粒度脉冲表示稀疏且离散,而全局表示是连续的,这使得传统的特征级融合易受异质表示间的干扰。因此,我们引入Sim-Fuse,一种相似度空间融合对齐策略,整合粗粒度与细粒度匹配关系,同时避免异质表示的直接融合。在Flickr30K和MSCOCO上的实验表明,我们的方法在匹配设置下优于人工神经网络(ANN)方法及现有的脉冲网络检索基线。此外,仅用2个时间步,我们的SNN即可达到与其ANN counterpart相当或更优的性能,同时将理论模块级能耗降低55%。代码在补充材料中提供。
英文摘要
Spiking neural networks (SNNs) offer an energy-efficient computing paradigm through sparse event-driven computation, showing great potential for efficient multimodal learning. However, applying SNNs to high-level multimodal tasks, such as image-text retrieval (ITR), remains challenging, since sparse spike representations make it difficult to capture semantic structures required for cross-modal alignment. Existing spiking ITR methods rely on local alignment and additional soft-label supervision during training, while lacking awareness of structural and multi-granularity relationships. To address these issues, we propose a Multi-head Spiking Graph Attention Network (\textbf{MSGAT}) for structural modeling and equip it with dynamic attention heads to capture complementary relational patterns and enable spike-driven graph reasoning and aggregation. However, within a two-branch multi-granularity fusion framework, the fine-grained spike representations generated by MSGAT are sparse and discrete, whereas the global representations are continuous, making conventional feature-level fusion susceptible to interference across heterogeneous representations. Therefore, we introduce \textbf{Sim-Fuse}, a similarity-space fusion alignment strategy integrating coarse- and fine-grained matching relations while avoiding direct fusion of heterogeneous representations. Experiments on Flickr30K and MSCOCO show our method outperforms ANN methods under matched settings and existing SNN retrieval baselines. Moreover, with only two time steps, our SNN achieves comparable or superior performance to its ANN counterpart while reducing theoretical module-level energy by 55\%. The code is provided in the Supplementary Materials.