arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向密集事件视觉的片上脉冲神经网络训练

Toward On-Chip Training of Spiking Neural Networks for Dense Event-Based Vision

Maxime Vaillant, Axel Carlier, Lai Xing Ng, Christophe Hurter, Benoit R. Cottereau

arXiv 2609.32405首次发表:更新:

AI 中文总结

针对事件相机密集视觉任务,提出DELL局部学习方案,以局部密集监督替代全局梯度传播,在DSEC光流上减少39.6%训练内存并提升精度,且参数少24倍仍具竞争力。

AI 中文摘要

事件相机为资源受限的机器人提供低延迟、异步的视觉感知。脉冲神经网络(SNN)天然地处理事件流,但使用时间反向传播(BPTT)训练深度SNN需要大量内存,并且在神经形态硬件上仍然困难。局部学习通过将误差传播限制在局部块内来避免这一问题,但现有方法主要针对分类而非密集预测。我们提出了DELL(密集事件驱动局部学习),一种用于密集事件视觉的块级方案,用局部密集监督替代全局梯度传播。可学习的、空间结构化的局部头在每个块适当的分辨率下进行监督,同时保留块内的时间动态。我们在光流回归和语义分割上使用全脉冲U形架构评估了DELL。在DSEC光流上,DELL相对于端到端BPTT将峰值训练内存减少了39.6%,同时提高了精度,在官方测试基准上达到1.670像素端点误差,而相同骨干网络端到端训练为1.941像素。块分离更像是一种正则化而非约束。现有的局部学习基线DECOLLE依赖固定的随机局部读出,不适合密集回归,导致端点误差高出3.9倍;可学习的局部头恢复了这一损失,并在所有光流指标上优于端到端训练。在分割任务上,它们恢复了大部分性能差距,尽管DELL仍比端到端训练低几个mIoU点。仅用2.3M参数,比最强SNN基线少24倍,该骨干网络在DSEC上仍与SNN最先进水平相当。这些结果将局部学习扩展到密集事件预测,同时大幅减少训练内存。

英文摘要

Event cameras provide low-latency, asynchronous visual sensing for resource-constrained robotics. Spiking neural networks (SNNs) process event streams naturally, but training deep SNNs with backpropagation through time (BPTT) requires substantial memory and remains difficult on neuromorphic hardware. Local learning avoids this by restricting error propagation to local blocks, but existing methods mainly target classification rather than dense prediction. We introduce DELL (Dense Event-driven Local Learning), a block wise scheme for dense event-based vision that replaces global gradient propagation with local dense supervision. Learnable, spatially structured local heads supervise each block at its appropriate resolution while preserving temporal dynamics within blocks. We evaluate DELL on optical-flow regression and semantic segmentation with a fully spiking U-shaped architecture. On DSEC optical flow, DELL reduces peak training memory by 39.6% relative to end-to-end BPTT while improving accuracy, reaching 1.670 px endpoint error on the official test benchmark versus 1.941 px for the same backbone trained end-to-end. Block detachment behaves more like a regularizer than a constraint. DECOLLE, the existing local-learning baseline, relies on fixed random local read-outs poorly suited to dense regression, resulting in a 3.9x higher endpoint error; learnable local heads recover this loss and outperform end-to-end training across all optical-flow metrics. On segmentation, they recover most of the performance gap, although DELL remains a few mIoU points behind end-to-end training. With 2.3M parameters, 24x fewer than the strongest SNN baseline, the backbone remains competitive with the SNN state of the art on DSEC. These results extend local learning to dense event-based prediction while substantially reducing training memory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑