SAGE:基于注意力引导熵的代理梯度适配脉冲 Transformer
SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers
- University of South Dakota(南达科他大学)
- USD Artificial Intelligence Research Lab(南达科他大学人工智能研究实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出 SAGE 机制,利用注意力熵估计的块级不确定性调整基于 Transformer 的脉冲神经网络的训练代理梯度斜率,在 CIFAR-10/100 上实现最高 1-2% 的稳定准确率提升,保留原始架构且部署成本不变。
AI中文摘要:
脉冲神经网络(SNN)通过利用稀疏的事件驱动计算,为传统深度神经网络提供了一种高能效的替代方案,但其训练仍面临挑战,因为不可微的脉冲函数需要代理梯度,而固定形状的代理梯度在不同层和训练阶段可能并非最优。在本研究中,我们提出 SAGE,一种用于基于 Transformer 的 SNN 的不确定性调制代理梯度机制。SAGE 从归一化自注意力熵中估计块级不确定性,并利用该信号在训练期间调整代理梯度的斜率,同时保持推理模型不变。仅在训练时调整代理参数,所提方法保留了原始架构和部署成本,同时提高了优化灵活性。在 CIFAR-10/100 上的实验表明,SAGE 相较于固定代理梯度基线实现了更高的准确率,在多个模拟时间步长上取得了最高达 1-2% 的稳定提升。这些结果凸显了源自注意力的不确定性作为轻量训练信号,在基于 Transformer 的 SNN 的自适应代理梯度学习中的潜力。
英文摘要:
Spiking neural networks (SNNs) offer an energy-efficient alternative to conventional deep neural networks by exploiting sparse event-driven computation, but their training remains challenging because the non-differentiable spike function requires surrogate gradients whose fixed shape may be suboptimal across layers and training stages. In this work, we introduce SAGE, an uncertainty-modulated surrogate-gradient mechanism for Transformer-based SNNs. SAGE estimates block-level uncertainty from normalized self-attention entropy and uses this signal to adapt the surrogate-gradient slope during training while leaving the inference model unchanged. By modulating only the training-time surrogate parameter, the proposed method preserves the original architecture and deployment cost while improving optimization flexibility. Experiments on CIFAR-10/100 demonstrate that SAGE achieves improved accuracy over fixed-surrogate baselines, with results up to 1-2\% consistent gains across multiple simulation time steps. These results highlight the potential of attention-derived uncertainty as a lightweight training signal for adaptive surrogate-gradient learning in transformer-based SNNs.