用于脉冲 Transformer 的脉冲局部交互与自适应互补融合
Spiking Local Interaction and Adaptive Complementary Fusion for Spiking Transformer
浏览论文内容
中文总结 AI 辅助
该研究针对脉冲 Transformer 局部空间上下文传播受限问题,提出脉冲局部交互与自适应互补融合方法,在多类视觉任务上实现性能提升,QKFormer 取得优异表现。
中文摘要 AI 辅助
脉冲 Transformer 主要通过脉冲自注意力(SSA)对 token 间的交互进行建模。然而,二值化的查询与键表示将连续相似度映射为稀疏离散的关系响应,可能会抑制弱关系并限制局部空间上下文的传播。为解决这一局限,我们提出脉冲局部交互(SLI)与自适应互补融合(ACF):SLI 利用轻量的深度-逐点变换,为相邻脉冲 token 间建立了一条不依赖注意力的直接信息交换通路;ACF 通过层特定的通道级系数,将 SSA 与 SLI 进行融合,自适应平衡二者在网络不同深度的贡献。该设计保留了原始注意力公式,可融入不同的脉冲 Transformer 架构,仅产生少量参数开销。在 ImageNet-1K、CIFAR-10、CIFAR-100、CIFAR10-DVS 和 ADE20K 上开展的实验表明,其在图像分类、事件驱动识别与语义分割任务中均取得了一致提升。具体而言,集成 SLI 与 ACF 的 QKFormer 在 ImageNet-1K 上达到 84.37% 的 Top-1 准确率,在未经过 ImageNet 预训练的情况下,于 ADE20K 上实现 37.5% 的 mIoU。消融研究与定性分析进一步显示,SSA 与 SLI 捕获互补的交互模式,且可学习的融合方式始终优于固定权重的方案。
英文摘要
Spiking Transformers model token interactions primarily through spiking self-attention (SSA). However, binary query and key representations map continuous similarities to sparse and discrete relation responses, which may suppress weak relations and limit the propagation of local spatial context. To address this limitation, we introduce Spiking Local Interaction (SLI) and Adaptive Complementary Fusion (ACF). SLI establishes an attention-independent pathway for direct information exchange among neighboring spiking tokens using lightweight depthwise--pointwise transformations. ACF integrates SSA and SLI through layer-specific, channel-wise coefficients that adaptively balance their contributions at different network depths. The proposed design preserves the original attention formulation and can be incorporated into different Spiking Transformer architectures with modest parameter overhead. Experiments on ImageNet-1K, CIFAR-10, CIFAR-100, CIFAR10-DVS, and ADE20K show consistent improvements across image classification, event-based recognition, and semantic segmentation. In particular, QKFormer with SLI and ACF achieves $84.37\%$ Top-1 accuracy on ImageNet-1K and $37.5\%$ mIoU on ADE20K, where the segmentation model is trained without ImageNet pretraining. Ablation studies and qualitative analyses further indicate that SSA and SLI capture complementary interaction patterns and that learnable fusion consistently outperforms fixed weighting.
发表机构
- Institute of Neuroscience, Chinese Academy of Sciences(中国科学院神经科学研究所)
- China Electric Power Research Institute Co., Ltd.(中国电力科学研究院有限公司)
机构由 AI 辅助整理,请以论文原文为准。