arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01622cs.NEcs.MM

SMM Transformer:利用脉冲神经网络实现多模态任务

SMM Transformer: Leveraging Spiking Neural Networks for Multimodal Tasks

Xiubo Liang, Jinxing Han, Yuke Li, Haoqi Zhu, Yu Zhao, Hongzhi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于脉冲神经网络的SMM Transformer框架,通过PLMP、SMSA、SMoE三个核心组件解决SNN多模态Transformer的训练与效率问题,在基准测试中实现了有竞争力的准确率与显著的能耗降低。

中文摘要 AI 辅助

脉冲神经网络(SNN)支持基于事件的计算,具有稀疏激活的特性,但在SNN上构建多模态Transformer面临两大阻碍:深度脉冲栈训练不稳定,以及密集softmax注意力与基于脉冲的通信不匹配。本文提出SMM Transformer,这是一种基于SNN的多模态Transformer框架,包含三个关键组件:一是PLMP,即并行LIF( leaky integrate-and-fire,泄漏积分放电)神经元,带有多级可学习参数,搭配定制的P-STBP算法以实现深度SNN的稳定训练;二是SMSA,一种受注意力启发的脉冲驱动令牌混合模块,用通道方向的脉冲共激活与自补偿替代密集的成对softmax注意力;三是SMoE,一种用于模态感知融合的脉冲混合专家模块。在视觉与多模态基准测试中,SMM Transformer与人工神经网络(ANN)基线相比达到了有竞争力的准确率。在标准MAC/AC算术模型下,SMSA将注意力模块的算子级计算能耗估计降低了最多97%,而全模型分析则显示出更温和但一致的效率提升。

英文摘要

Spiking Neural Networks (SNNs) enable event-driven computation with sparse activations, but building multimodal Transformers on SNNs is hindered by unstable training in deep spiking stacks and the mismatch between dense softmax attention and spike-based communication. We propose SMM Transformer, an SNN-based multimodal Transformer framework that combines (i)PLMP, a Parallel LIF with Multistage Learnable Parameters neuron and a tailored P-STBP algorithm for stable deep SNN training, (ii) SMSA, an attention-inspired spike-driven token-mixing module that replaces dense pairwise softmax attention with channel-wise spike co-activation and self-compensation, and (iii)SMoE, a spiking mixture-of-experts module for modality-aware fusion. Across visual and multimodal benchmarks, SMM Transformer achieves competitive accuracy compared to ANN baselines. Under a standard MAC/AC arithmetic model, SMSA reduces the estimated operator-level compute energy of the attention module by up to 97%, while whole-model profiling shows more moderate but consistent efficiency gains.

↑