隐藏于明处:针对视觉-语言-动作模型的基于扩散的无约束机器人攻击
Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models
AI总结:
本研究提出基于扩散的无约束机器人攻击方法DURA,针对视觉-语言-动作模型生成自然对抗补丁,实验表明其性能优于现有方法,暴露了物理部署的VLA模型的安全风险。
AI中文摘要:
视觉-语言-动作(VLA)模型在控制机器人完成各类操纵任务方面展现出强大能力,但其对抗鲁棒性仍未得到充分探索,利用这一弱点可能导致现实世界的危害。现有针对VLA模型的攻击通常依赖像素空间扰动或白盒访问,会产生明显伪影,在现实机器人系统中的部署能力有限。本研究提出DURA,一种基于扩散的无约束机器人攻击,可为VLA模型生成视觉自然的对抗补丁。DURA支持白盒和黑盒攻击设置,其中黑盒设置仅需目标模型的预测动作。通过在预训练扩散模型的潜在轨迹上优化,DURA生成视觉自然的补丁,同时引导机器人执行攻击者指定的目标动作。在仿真和真实物理世界中的大量实验表明,DURA的性能始终优于现有方法。我们的发现揭示了物理部署的VLA模型存在安全风险,呼吁构建更强的防御机制。
英文摘要:
Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial robustness remains largely underexplored, and exploiting this weakness can lead to physical-world harm. Existing attacks on VLA models often rely on pixel-space perturbations or white-box access, resulting in noticeable artifacts and limited deployability in real-world robotic systems. In this work, we propose DURA, a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches for VLA models. DURA supports both white-box and black-box attack settings, where the black-box setting requires only the predicted actions of the victim model. By optimizing along the latent trajectory of a pretrained diffusion model, DURA generates visually natural patches while steering the robot toward attacker-specified target actions. Extensive experiments in both simulation and the real physical world show that DURA consistently outperforms existing methods. Our findings expose a safety risk for physically deployed VLA models and call for stronger defenses.