发表机构
Guangdong Institute of Intelligence Science and Technology; Hong Kong Polytechnic University; The University of Hong Kong; University of Illinois Urbana-Champaign(广东智能科学与技术研究院; 香港理工大学; 香港大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DS-VLA通过树突脉冲动力学和神经元级抑制门增强VLA模型鲁棒性,在LIBERO基准上扰动下保持95.4%性能,优于现有方法。
AI 中文摘要
视觉-语言-动作(VLA)模型在语言条件化操作任务中已展现出强劲性能,然而,在名义评估中的成功并不一定意味着当执行动作受到瞬时扰动时,模型仍能表现出鲁棒的闭环行为。我们提出了DS-VLA,一种树突启发的动作架构,将树突脉冲动力学融入VLA控制以解决上述局限。具体而言,为实现模块化特征处理和时间信息整合,DS-VLA为动作神经元配备了多个稀疏连接的树突分支,每个分支具有异质的学习衰减因子。此外,为抑制不可靠的状态更新同时保留任务相关的历史信息,我们引入了一种神经元级抑制门,在体细胞动力学之前自适应地调节多模态新证据进入树突状态的准入。我们在全部四个LIBERO套件上,在名义轨迹和统一的闭环动作扰动协议下评估了DS-VLA。DS-VLA实现了91.6%的平均名义成功率和87.35%的平均扰动成功率,保留了其名义性能的95.4%。在相同的报告扰动设置下,OpenVLA-OFT、FAST、π0和GR00T分别实现了39.45%、23.90%、28.55%和30.75%的成功率。一项受控消融实验分离了神经元级共享抑制的贡献,而神经动力学和扰动后轨迹的分析将鲁棒性能与选择性证据抑制和有效行为恢复相关联。这些结果共同表明,整合大脑启发的计算机制为鲁棒具身智能提供了一种有前景的架构先验,而不仅仅是扩展视觉-语言骨干网络或生成式动作解码器。
英文摘要
Vision-language-action (VLA) models have achieved strong performance in language-conditioned manipulation, yet success under nominal evaluation does not necessarily translate into robust closed-loop behavior when executed actions are transiently corrupted. We introduce DS-VLA, a dendritic-inspired action architecture that incorporates dendritic spiking dynamics into VLA control to address this limitation. Specifically, to enable modularized feature processing and temporal information integration, DS-VLA equips action neurons with multiple sparsely connected dendritic branches, each featuring heterogeneous, learned decay factors. Furthermore, to suppress unreliable state updates while preserving task-relevant historical information, we introduce a neuron-wise inhibitory gate that adaptively regulates the admission of new multimodal evidence into dendritic states prior to somatic dynamics. We evaluate DS-VLA on all four LIBERO suites under both nominal rollouts and a unified closed-loop action-perturbation protocol. DS-VLA achieves a 91.6\% average nominal success rate and an 87.35\% average perturbed success rate, retaining 95.4\% of its nominal performance. Under the same reported perturbation setting, OpenVLA-OFT, FAST, $π_0$, and GR00T achieve 39.45\%, 23.90\%, 28.55\%, and 30.75\%, respectively. A controlled ablation isolates the contribution of neuron-wise shared inhibition, while analyses of neural dynamics and post-perturbation trajectories associate robust performance with selective evidence suppression and effective behavioral recovery. Together, these results demonstrate that integrating brain-inspired computational mechanisms offers a promising architectural prior for robust embodied intelligence beyond merely scaling vision-language backbones or generative action decoders.