arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07066cs.AI

PTQ4SNN:面向脉冲神经网络的感知膜后训练量化

PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks

Hui Xie, Tong Shi, Haotong Qin, Aishan Liu, Xiaode Liu, Jinyang Guo

首次发表
浏览论文内容

中文总结 AI 辅助

PTQ4SNN是一种仅用小型校准集联合量化SNN权重与循环膜状态的感知膜后训练量化框架,可在无需重训骨干网络的情况下,于W4量化和约4位膜精度下保持模型准确率。

中文摘要 AI 辅助

脉冲神经网络(SNN)支持稀疏且基于事件的计算,但其低位部署尚未完善,因为即使在权重量化后,循环膜状态通常仍以浮点数形式保留。对这些状态进行量化颇具挑战,原因在于其分布因通道而异,且与先前的权重分布不同,同时接近发放阈值的微小扰动可能会改变脉冲决策并随时间累积。我们提出PTQ4SNN,这是一种感知膜的后训练量化框架,仅使用小型校准集即可联合量化权重和循环膜状态。首先,一种通道级统一尺度桥将膜尺度约束为s_mem,c = s_w,c * 2^k_c,以适应膜分布,同时实现与移位兼容的尺度转换。其次,混合精度比特分配在平均比特预算下,根据发放活动和量化敏感度为膜通道分配2/4/8位精度。该框架适用于可重用的投影-LIF对,支持卷积SNN和脉冲驱动的Transformer,无需对骨干网络进行重新训练。在静态与基于事件的分类以及语义分割任务上的实验表明,PTQ4SNN在W4量化和约4位膜精度下可有效保持模型准确率。

英文摘要

Spiking neural networks (SNNs) enable sparse and event-driven computation, but their low-bit deployment remains incomplete because recurrent membrane states are commonly retained in floating point even after weight quantization. Quantizing these states is challenging because their distributions differ across channels and from the preceding weights, while small perturbations near the firing threshold may alter spike decisions and accumulate over time. We propose PTQ4SNN, a membrane-aware post-training quantization framework that jointly quantizes weights and recurrent membrane states using only a small calibration set. First, a channel-wise Unified Scale Bridge constrains the membrane scale as s_mem,c = s_w,c * 2^k_c, adapting to membrane distributions while enabling shift-compatible scale conversion. Second, Mixed-Precision Bit Allocation assigns 2/4/8-bit precision to membrane channels according to firing activity and quantization sensitivity under an average-bit budget. The framework operates on reusable projection-LIF pairs and supports both convolutional SNNs and spike-driven Transformers without backbone retraining. Experiments on static and event-based classification and semantic segmentation show that PTQ4SNN effectively preserves model accuracy under W4 quantization and approximately 4-bit membrane precision.

↑