arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边推理边学习:面向跨感知模态边缘SNN的局部并行学习

Learning While Inferring: Local and Parallel Learning for Edge SNNs across Sensing Modalities

Yanxun Zhang, Yifei Wang, Changze Lv, Jingwen Xu, Yiyang Lu, Xiaohua Wang, Di Yu, Xin Du, Xiaoqing Zheng

arXiv 2610.03149首次发表:更新:

发表机构

College of Computer Science and Artificial Intelligence, Fudan University; Shanghai Key Laboratory of Intelligent Information Processing; School of Software Technology, Zhejiang University(复旦大学计算机科学技术学院; 上海市智能信息处理重点实验室; 浙江大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出双向脉冲蒸馏(BSD)作为边缘SNN的在线学习原则,通过局部目标对齐前向与反向网络,实现边推理边学习,在25个基准上接近BP性能,且训练延迟和能耗大幅降低。

AI 中文摘要

边缘智能要求模型在严格的算力、能耗和内存预算下,实时连续感知并持续在设备端适应。尽管脉冲神经网络(SNN)支持高效的事件驱动推理,但标准的替代梯度反向传播(BP)会串行化更新并阻塞进行中的推理。我们研究双向脉冲蒸馏(BSD)作为一种设备端学习原则,使边缘SNN能够边推理边学习。BSD将刺激驱动的前向网络与独立的目标驱动反向网络耦合,并通过局部目标对齐它们的中间表示。由于两条通路在对齐之前具有不相交的计算图,前向推理、反向推理和分阶段更新可以并发运行。在涵盖SOUL五种感知模态的25个基准上,BSD平均保持在匹配的BP基线3.8个百分点以内,而部署时仅需其前向分支。学习到的表示也能很好地迁移到少样本类增量学习,无需重放过去的数据。通过移除密集的浮点反向链,BSD将预计训练延迟降至BP的0.72倍,预计训练能耗降至BP的0.36倍。

英文摘要

Edge intelligence requires models to sense continuously in real time and to keep adapting on-device, all under tight compute, energy, and memory budgets. Although spiking neural networks (SNNs) enable efficient event-driven inference, standard surrogate-gradient backpropagation (BP) serializes updates and blocks ongoing inference. We investigate Bidirectional Spike-Based Distillation (BSD) as an on-device learning principle that lets edge SNNs learn while inferring. BSD couples a stimulus-driven forward network with an independent target-driven reverse network and aligns their intermediate representations through local objectives. Because the two pathways have disjoint computation graphs until alignment, forward inference, reverse inference, and stage-wise updates can run concurrently. On 25 benchmarks spanning the five sensing modalities of SOUL, BSD stays within 3.8 percentage points of matched BP baselines on average, while only its forward branch is needed at deployment. The learned representations also transfer well to few-shot class-incremental learning without replaying past data. By removing the dense floating-point backward chain, BSD reduces the projected training latency to $0.72\times$ and the estimated training energy to $0.36\times$ that of BP.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑