发表机构
School of Computer Science, Wuhan University; D-Robotics; HUST; Institute of Automation, Chinese Academy of Sciences(武汉大学计算机科学学院; D-机器人公司; 华中科技大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉-语言-动作模型中毒后门问题,通过两种攻击识别出紧凑因果足迹机制,基于此提出TrustVLA防御,能检测异常证据演变、定位支持并恢复观测,在多评估中降低攻击成功率且保持干净任务性能。
AI 中文摘要
视觉-语言-动作(VLA)模型通过终端用户无法审计的管道部署,中毒的VLA在干净观测上表现正常,小视觉触发器会在故障可观测前重定向长期机器人策略。现有视觉或语言防御很少解释触发的VLA表示形式或如何在不重新训练的情况下恢复行为。我们通过BadVLA和INFUSE这两种具有不同注入策略的VLA攻击研究这一差距。我们识别出一种反复出现的内部机制——紧凑因果足迹,基于此提出TrustVLA防御,它能检测异常证据演变、定位紧凑支持并通过局部修复恢复观测。在跨OpenVLA/LIBERO和$\pi_{0.5}$转移评估中,TrustVLA降低攻击成功率并保持干净任务性能。
英文摘要
Vision-Language-Action (VLA) models are deployed through pipelines that end users cannot audit, and a poisoned VLA can behave normally on clean observations while a small visual trigger redirects a long-horizon robot policy before any failure becomes observable. Existing vision or language defenses rarely explain what a triggered VLA representation looks like or how to recover behavior without retraining. We study this gap through two independently proposed VLA attacks from groups with distinct injection strategies, BadVLA and INFUSE; the latter persists after downstream clean adaptation. Across the evaluated poisoned models, we identify a recurring internal mechanism: a \emph{compact causal footprint}, namely a small visual support that is attention-seeded, spatially compact, and \emph{causal} in a precise sense -- masking it returns a clean-calibrated evidence-evolution score to the normal operating region. This footprint motivates TrustVLA, a mechanism-guided inference-time defense that adapts the Dirichlet evidence framework from trusted classification to monitor per-token, per-layer epistemic uncertainty in VLA policies. With only a small clean calibration set, TrustVLA (i)~detects abnormal evidence evolution, (ii)~localizes the compact support by counterfactual mechanism-score drop, and (iii)~recovers the observation by localized inpainting. Across OpenVLA/LIBERO and $π_{0.5}$ transfer evaluations, TrustVLA reduces attack success while preserving clean-task performance, providing a retraining-free, mechanism-guided defense for visual-triggered VLA backdoors.