CertVLA:针对视觉-语言-动作模型的物理视觉攻击的可验证防御
CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models
AI总结:
本文提出CertVLA,一种针对VLA模型的物理视觉攻击的可验证防御,通过校准动作区域、确定性覆盖掩码等技术,可验证闭环控制的动作安全性,经仿真和真实实验验证其有效性。
AI中文摘要:
视觉-语言-动作(VLA)策略易受局部物理扰动影响,但现有可验证补丁防御针对离散标签,无法直接验证连续、时间相关的动作。本文提出CertVLA,这是一种针对有界补丁和纹理攻击的闭环VLA控制的可验证防御方法。CertVLA提出了行为一致动作的校准区域,同时确定性覆盖掩码确保至少一个经过检查的预测无攻击。具体而言,CertVLA通过每对掩码的良性变化归一化动作分歧,仅当单掩码锚点在每第二个掩码下保持一致时才接受它;随后校准所得的max-min-max回合分数以提供有限样本的干净覆盖。结合查询级决策将动作证书扩展到完整的闭环滚动。此外,我们证明,针对满足有界支持威胁模型的任何自适应攻击者,CertVLA验证的每个滚动仅执行与攻击消除后的干净预测一致的动作块。在双掩码滚动正确性下,此一致性证书进一步保证任务成功。该证书独立于补丁内容、生成方法和物理变换。仿真和真实世界实验证明了CertVLA针对补丁攻击的经验和可验证有效性,还在仿真中对纹理攻击进行了额外验证。
英文摘要:
Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target discrete labels and cannot directly certify continuous, temporally correlated actions. We introduce CertVLA, a certified defense for closed-loop VLA control under bounded patch and texture attacks. CertVLA proposes a calibrated region of behaviorally consistent actions, while deterministic covering masks ensure that at least one checked prediction is attack-free. Specifically, CertVLA normalizes action disagreement by the benign variation of each mask pair and accepts a single-mask anchor only when it remains consistent under every second mask. It then calibrates the resulting max-min-max episode score to provide finite-sample clean coverage. Conjoining query-level decisions extends the action certificate to the complete closed-loop rollout. Furthermore, we prove that against any adaptive attacker satisfying the bounded-support threat model, every rollout certified by CertVLA executes only action chunks consistent with attack-erased clean predictions. Under dual-mask rollout correctness, this consistency certificate further guarantees task success. The certificate is independent of patch content, generation method, and physical transformation. Experiments in simulation and the real world demonstrate the empirical and certified effectiveness of CertVLA against patch attacks, with additional simulation validation on texture attacks.