脉冲驱动的视觉-语言-动作模型
Spike-driven Vision-Language-Action Model
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出首个脉冲驱动的VLA框架,通过脉冲编码器、多胜者融合和脉冲动作分块Transformer,实现低能耗、端到端的机器人操作,在LIBERO和Meta-World上以更少参数达到竞争性能。
AI中文摘要:
视觉-语言-动作(VLA)模型弥合了多模态理解与机器人控制之间的鸿沟,推动了具身智能的主导范式。然而,现有大多数模型依赖大型Transformer,其延迟和能耗成本阻碍了在资源受限平台上的部署。通过稀疏事件驱动计算,脉冲神经网络为高性能和节能计算提供了一种有前景的范式。在此,我们提出了首个脉冲驱动的VLA框架,支持机器人操作任务的端到端直接训练,该框架主要由三个核心组件构成。首先,我们开发了用于多模态感知的脉冲视觉和指令编码器,将视觉观察和语言指令编码为稀疏、可靠的脉冲表示,以供后续跨模态融合使用。其次,我们引入了多胜者脉冲融合机制用于指令引导的场景理解,通过双向top-$k$胜者全取脉冲路由来抑制背景干扰并生成融合记忆。最后,我们提出了一种脉冲动作分块Transformer,该模型在融合记忆和当前机器人状态上应用脉冲交叉注意力,从而高效地端到端生成连续动作块以进行机器人控制。在LIBERO和Meta-World上的大量实验表明,脉冲驱动的VLA在参数更少、估计推理能耗更低的情况下,达到了与常规VLA模型相当的性能。这项工作为神经形态VLA建模建立了基础框架,为未来资源高效的具身智能发展铺平了道路。
英文摘要:
Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. However, most existing models rely on large Transformers, whose latency and energy costs hinder deployment on resource-constrained platforms. Through sparse event-driven computation, spiking neural networks offer a promising paradigm for high-performance and energy-efficient computing. Here, we propose the first Spike-driven VLA framework enabling end-to-end direct training for robotic manipulation, which mainly comprises three core components. First, we develop spiking visual and instruction encoders for multimodal perception, encoding visual observations and language instructions into sparse, reliable spike representations for subsequent cross-modal fusion. Then, we introduce Multi-Winner Spike Fusion for instruction-guided scene understanding, using bidirectional top-$k$ winner-take-all spike routing to suppress background interference and yield fused memory. Finally, we propose a Spike Action Chunking Transformer that incorporates spiking cross-attention over the fused memory and the current robot state, enabling efficient end-to-end generation of continuous action chunks for robotic control. Extensive experiments on LIBERO and Meta-World demonstrate that Spike-driven VLA achieves competitive performance with fewer parameters and lower estimated inference energy than conventional VLA models. This work establishes a foundational framework for neuromorphic VLA modeling, paving the way for future advances in resource-efficient embodied intelligence.