发表机构
University of Electronic Science and Technology of China; Shenzhen Loop Area Institute(电子科技大学; 深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文系统分析内在位置学习(IPL)在事件驱动脉冲跟踪中的噪声问题,提出计算图裁剪方法Mask IPL,通过有效性掩码消除噪声,在不增参数下提升跟踪性能。
AI 中文摘要
脉冲神经网络(SNNs)与事件相机的事件驱动特性相匹配,并能自然地提取时空特征。这些特性促使了一系列近期关于基于SNNs的事件跟踪研究。内在位置学习(IPL)在不引入额外参数的情况下获取强位置信息,使其成为事件驱动脉冲跟踪中位置编码的主流方法。然而,其有效性背后的机制缺乏系统的理论分析。此外,我们的分析揭示,IPL在前向和反向传播中均引入噪声。前者增加推理误差,而后者阻止参数收敛到更优解。本文对IPL进行了系统分析,并证明其有效性源于IPL与多阶段卷积之间的协同作用。联合张量中的零块充当卷积的零填充,由此产生的边界效应通过多阶段卷积逐层传播。因此,每次参数更新都由一个能感知模板帧与搜索帧之间相对位移的梯度驱动。在卷积阶段之后添加的位置编码无法提供此信息。我们进一步提出一种简单的计算图裁剪方法,该方法根据布局确定的有效性掩码应用于每一层的操作,使无效区域在前向和反向传播中均等同于零填充。这在不引入额外参数的情况下消除了噪声,并使实际梯度与理想梯度一致。我们将改进的方法命名为Mask IPL。在不增加参数或计算成本的情况下,Mask IPL在FE108、FELT和VisEvent上提高了Tiny规模跟踪器的AUC,并持续改进Base规模跟踪器。
英文摘要
Spiking Neural Networks (SNNs) match the event-driven nature of event cameras and naturally extract spatiotemporal features. These properties have motivated a series of recent studies on event-based tracking with SNNs. Intrinsic Position Learning (IPL) acquires strong position information without introducing additional parameters, making it a mainstream approach for position encoding in event-based spike-driven tracking. However, the mechanism behind its effectiveness lacks systematic theoretical analysis. Moreover, our analysis reveals that IPL introduces noise in both forward and backward propagation. The former increases inference error, while the latter prevents parameters from converging to better solutions. This paper presents a systematic analysis of IPL and demonstrates that its effectiveness stems from the synergy between IPL and multi-stage convolution. The zero blocks in the joint tensor act as zero padding for convolution, and the resulting boundary effect propagates layer by layer through multi-stage convolution. Every parameter update is therefore driven by a gradient that perceives the relative displacement between template and search frames. Positional encoding added after the convolutional stage cannot provide this information. We further propose a simple Computation Graph Clipping method that applies a validity mask determined by the layout to the operations of every layer, making invalid regions equivalent to zero padding in both forward and backward propagation. This eliminates the noise without introducing additional parameters and makes the actual gradient coincide with the ideal gradient. We name the improved method Mask IPL. Without increasing parameters or computational cost, Mask IPL improves the AUC of the Tiny-scale tracker on FE108, FELT, and VisEvent, and consistently improves the Base-scale tracker as well.