arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考脉冲Transformer中的注意力局部性

Rethinking Attention Locality in Spiking Transformers

Zeqi Zheng, Zizheng Zhu, Yuping Yan, Wenxuan Pan, Zhaofei Yu, Yaochu Jin

arXiv 2608.08541首次发表:更新:

发表机构

Zhejiang University; Westlake University; Peking University(浙江大学; 西湖大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对脉冲Transformer中注意力局部性的问题,提出带边界连续性通路的空间连续局部注意力(SCLA-BCP)及分层部署策略,在7个静态和神经形态数据集上实现了检测、分割等任务的性能提升。

AI 中文摘要

脉冲Transformer为基于脉冲驱动计算的高效视觉处理提供了一种有前景的范式,但其无Softmax的脉冲自注意力(SSA)难以建立空间上局部的令牌交互。尽管现有的局部性增强SSA方法提升了准确率,但仍不清楚它们是否能在各层及不同脉冲Transformer架构中一致地诱导空间局部性。通过平均注意力距离(MAD)分析,我们发现计算局部性不一定转化为空间局部性,且统一应用相同的局部性增强会忽略架构依赖的部署需求。基于这些观察,我们提出了带边界连续性通路的空间连续局部注意力(SCLA-BCP)。SCLA在空间相邻令牌的非重叠区域内计算注意力,而BCP则通过轻量卷积通路促进跨边界信息交换。此外,我们开发了分层局部性部署策略,以将SCLA-BCP有效应用于两种主要的脉冲Transformer架构。在涵盖分类、检测和分割任务的7个静态和神经形态数据集上进行的大量实验表明,该方法在参数和能量开销有限的情况下取得了一致的性能提升。值得注意的是,我们的方法在COCO 2017上将mAP@50提升了最多9.50%,在ADE20K上将mIoU提升了最多3.42%。可视化、MAD分析和 ablation 研究进一步验证了其有效性。

英文摘要

Spiking Transformers provide a promising paradigm for efficient visual processing with spike-driven computation, yet their Softmax-free Spiking Self-Attention (SSA) struggles to establish spatially localized token interactions. Although existing locality-enhanced SSA methods improve accuracy, it remains unclear whether they consistently induce spatial locality across layers and different Spiking Transformer architectures. Through Mean Attention Distance (MAD) analysis, we reveal that computational locality does not necessarily translate into spatial locality and show that uniformly applying the same locality enhancement overlooks architecture-dependent deployment requirements. Motivated by these observations, we propose Spatially Contiguous Local Attention with Boundary Continuity Pathway (SCLA-BCP). SCLA computes attention within non-overlapping regions of spatially adjacent tokens, while BCP facilitates cross-boundary information exchange through a lightweight convolutional pathway. Furthermore, we develop a hierarchical locality deployment strategy to effectively apply SCLA-BCP across the two major Spiking Transformer architectures. Extensive experiments on seven static and neuromorphic datasets covering classification, detection, and segmentation demonstrate consistent improvements with limited parameter and energy overhead. Notably, our approach improves mAP@50 by up to 9.50% on COCO 2017 and mIoU by up to 3.42% on ADE20K. Visualizations, MAD analysis, and ablation studies further validate its effectiveness.

Comments17 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑