发表机构
The Hong Kong University of Science and Technology; Beihang University(香港科技大学; 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出利用VLA祖先VLM的黑盒对抗补丁攻击,通过破坏继承的视觉感知和指令接地能力,显著降低任务成功率,并分析脆弱性继承的边界条件。
AI 中文摘要
视觉-语言-动作模型(VLAs)正越来越多地部署在安全关键的物理环境中,然而其对抗鲁棒性仍鲜为人知。现有攻击通常假设白盒访问或依赖替代VLA,这在现实部署中很少成立。我们的关键洞察是,大多数VLA是从公开发布的预训练视觉语言模型(VLM)适配而来,继承了动作生成所必需的两个能力:视觉感知和指令条件接地。因此,本文探讨了一个先前未解决的问题:攻击者能否仅使用其祖先VLM来攻击已部署的VLA?为此,我们提出了三种对抗补丁攻击,以破坏这些继承的能力:一种通过相对和绝对项破坏投影视觉令牌的视觉干扰攻击,一种移除指令接地概念所需视觉证据的指令接地语义证据抑制攻击,以及一种在双阶段课程下统一两个目标的联合攻击。在模拟和静态真实世界图像上的不同VLA家族实验中表明,在祖先VLM上优化的补丁导致VLA任务成功率显著下降,证明VLA在继承基础能力的同时也继承了对抗脆弱性。这种效应并非均匀:在需要精确指令接地定位的任务上最强,而在适配重写了共享视觉语义表示或动作头迭代平滑扰动策略上几乎消失。通过刻画脆弱性继承的边界条件,并提供继承效应成立或失败原因的分析,我们推进了对涉及VLA系统安全性的理解。
英文摘要
Vision-Language-Action models (VLAs) are increasingly deployed in safety-critical physical environments, yet their adversarial robustness remains poorly understood. Existing attacks typically assume white-box access or rely on surrogate VLAs, which rarely holds in real-world deployments. Our key insight is that most VLAs are adapted from a publicly released pretrained vision-language model (VLM), inheriting two capabilities essential for action generation: visual perception and instruction-conditioned grounding. Therefore, this paper explores a previously unaddressed question: can an adversary attack deployed VLAs using only their ancestor VLMs? To this end, we propose three adversarial patch attacks that disrupt the inherited capabilities: a vision disruption attack that corrupts the projected visual tokens through relative and absolute terms, an instruction-grounded semantic evidence suppression attack that removes the visual evidence required for instruction-grounded concepts, and a joint attack that unifies both objectives under a two-phase curriculum. Experiments across different VLA families on both simulation and static real-world images show that patches optimized on the ancestor VLM cause substantial degradations in VLA task success rates, demonstrating that VLAs inherit adversarial vulnerabilities alongside their foundational capabilities. This effect is not uniform: it is strongest on tasks that require precise instruction-grounded localization, and nearly vanishes on policies whose adaptation rewrites the shared visual-semantic representation or whose action head iteratively smooths perturbations away. By characterizing the boundary conditions of vulnerability inheritance and providing analysis of why the inheritance effect holds or fails, we advance the understanding of safety for VLA-involved systems.