发表机构
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ); Shenzhen University(广东省人工智能与数字经济实验室(深圳); 深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VLAGuard框架通过提出VASA攻击和APFT防御,有效降低了无线传感器网络中VLA机器人的物理注意力劫持风险,提升了其在模拟与真实环境中的鲁棒性。
AI 中文摘要
将视觉-语言-动作(VLA)机器人作为移动边缘节点部署在无线传感器网络(WSNs)中,需要具备抵御物理对抗性威胁的强大防护能力。本文提出VLAGuard框架,用于评估和缓解一种关键漏洞:策略关键型动作到视觉的注意力劫持。我们首先引入压力测试模块——视觉运动注意力引导语义攻击(VASA),该模块使用可打印补丁严重干扰机器人的动作条件交叉注意力。为应对这一问题,我们提出注意力保护微调(APFT)防御方法,该方法可稳定时空注意力并强制几何一致性,且无推理开销。在模拟和物理WSN辅助的智能环境中的评估显示,鲁棒性显著提升:在LIBERO模拟中,APFT将OpenVLA的失败率从100.0%降至25.9%;此外,在2000次真实试验中,APFT在严重补丁攻击下将平均成功率从23.0%提升至67.4%。这表明保护注意力通路对提升传感器网络中VLA驱动边缘节点的鲁棒性具有重要意义。
英文摘要
Deploying Vision-Language-Action (VLA) robots as mobile edge nodes within wireless sensor networks (WSNs) requires robust protection against physical adversarial threats. We present VLAGuard, a framework to assess and mitigate a critical vulnerability: policy-critical action-to-vision attention hijacking. We first introduce a stress-test module, Visuomotor Attention-guided Semantic Attack (VASA), using printable patches to severely distract the robot's action-conditioned cross-attention. To counter this, we propose Attention-Protective Fine-Tuning (APFT), a defense that stabilizes spatiotemporal attention and enforces geometric consistency with zero inference overhead. Evaluations across simulated and physical WSN-assisted smart environments demonstrate significant robustness gains. APFT reduces the OpenVLA failure rate from 100.0% to 25.9% in LIBERO simulations. Furthermore, across 2,000 real-world trials, APFT improves the average success rate from 23.0% to 67.4% under severe patch attacks. This highlights that protecting attention pathways is important for improving the robustness of VLA-driven edge nodes in sensor networks.
Comments32 pages, 8 figures, 5 tables. Accepted for publication in Ad Hoc & Sensor Wireless Networks (AHSWN)