arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11865cs.NE

Lapis:基于首次脉冲时序与膜泄漏的拉普拉斯脉冲注意力机制

Lapis: Laplacian Spiking Attention via First-Spike Timing and Membrane Leakage

Kaiwen Tang, Jiaqi Zheng, Zixuan Zhu, Yiqun Wang, Zhanglu Yan, Weng-Fai Wong

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出基于首次脉冲时序与膜泄漏的Lapis脉冲注意力机制,仅用加减等操作实现低能耗评分,在CIFAR-10和ImageNet-1K上取得接近点积注意力的准确率,大幅降低了注意力路径的算术能耗。

中文摘要 AI 辅助

自注意力已成为脉冲视觉Transformer的核心,但其查询-键评分仍在很大程度上继承自密集网络。现有脉冲变体要么简化了点积评分,要么将其替换为离散算子,但脉冲时序作为脉冲网络的固有变量,并未直接定义令牌间的关联方式。我们提出Lapis,一种脉冲注意力机制,其通过时间至首次脉冲编码下每个令牌对的查询与键首次脉冲延迟向量的L1距离对令牌对进行评分,并通过拉普拉斯核将该距离映射为亲和度。该核的指数衰减与漏极整合-发放膜的脉冲响应相匹配,因此累积的延迟差决定了膜迹线的衰减,而行归一化在二的幂次取整下简化为移位操作。因此,评分仅需减法、绝对值与累加操作,且消除了查询与键通道间的所有乘法操作。在匹配的骨干网络与训练计划下,Lapis在CIFAR-10上达到96.56%的Top-1准确率,与点积评分的差距在0.53个百分点以内;在ImageNet-1K上,相比密集点积注意力,其注意力路径的估计算术能耗降低了14.5倍,部署的6位模型在每张图像估计算术能耗为3.28mJ时达到83.25%的Top-1准确率。

英文摘要

Self-attention has become central to spiking vision transformers, yet its query-key scoring is still largely inherited from dense networks. Existing spiking variants either simplify dot product scoring or replace it with discrete operators, but spike timing, the native variable of a spiking network, does not directly define how tokens are related. We propose Lapis, a spiking attention mechanism that scores each token pair by the L1 distance between its query and key first-spike latency vectors under time-to-first-spike coding, and maps this distance to an affinity through a Laplacian kernel. The kernel's exponential decay matches the impulse response of a leaky integrate-and-fire membrane, so the accumulated latency difference determines the decay of a membrane trace, while row normalization reduces to a bit shift under power-of-two rounding. Scoring therefore needs only subtraction, absolute value, and accumulation, and removes all multiplication between query and key channels. Under a matched backbone and training schedule, Lapis reaches 96.56% top-1 accuracy on CIFAR-10, within 0.53 points of dot-product scoring. On ImageNet-1K, it reduces the estimated arithmetic energy of the attention path by 14.5x relative to dense dot-product attention. The deployed 6-bit model attains 83.25% top-1 accuracy at an estimated arithmetic energy of 3.28mJ per image.

补充信息

↑