发表机构
University of Electronic Science Technology of China; Shenzhen Loop Area Institute; The Chinese University of Hong Kong (Shenzhen)(电子科技大学; 深圳河套学院; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SpikingVLA通过树突整合-发放神经元和异步执行机制,实现低延迟、高能效的尖峰VLA推理,显著提升导航性能与效率。
AI 中文摘要
ANN到SNN的转换提供了一条实用的途径,通过绕过从头训练大规模SNN的巨额成本,实现能效高的尖峰视觉-语言-动作(VLA)模型。然而,现有方法通常需要许多时间步来维持有竞争力的性能,导致实时VLA部署产生大量推理延迟。为解决这一挑战,我们提出了SpikingVLA,一个ANN到SNN的转换框架,能够实现准确且低延迟的尖峰VLA推理。具体而言,我们提出了一种树突整合-发放(DIF)神经元,通过树突混合和自适应胞体发放来缓解通道级激活异常,从而以更少的时间步实现准确的ANN到SNN转换。基于DIF神经元,我们进一步引入了一种异步执行机制,该机制在VLA组件之间重叠时间计算,减少同步开销和延迟。大量实验表明,SpikingVLA在显著提高推理效率的同时,实现了有竞争力的导航性能。与现有的尖峰VLA方法相比,SpikingVLA将SR和SPL分别提高了11.9%和12.6%,同时将首次动作延迟降低了11.2倍。这些结果确立了SpikingVLA作为一个实用框架,用于以高性能和低延迟的尖峰推理部署预训练的VLA模型。
英文摘要
ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference latency for real-time VLA deployment. To address this challenge, we introduce SpikingVLA, an ANN-to-SNN conversion framework that enables accurate and low-latency spiking VLA inference. Specifically, we propose a Dendritic Integrate-and-Fire (DIF) neuron that alleviates channel-wise activation outliers through dendritic mixing and adaptive somatic firing, enabling accurate ANN-to-SNN conversion with fewer timesteps. Building on DIF neurons, we further introduce an asynchronous execution mechanism that overlaps temporal computation across VLA components, reducing synchronization overhead and latency. Extensive experiments demonstrate that SpikingVLA achieves competitive navigation performance with substantially improved inference efficiency. Compared with existing spiking VLA methods, SpikingVLA improves SR and SPL by 11.9\% and 12.6\%, respectively, while reducing first-action latency by 11.2$\times$. These results establish SpikingVLA as a practical framework for deploying pretrained VLA models with high-performance and low-latency spiking inference.