DRIFT:通过对抗性补丁攻击破坏流匹配视觉-语言-动作(VLA)模型的去噪轨迹
DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack
浏览论文内容
中文总结 AI 辅助
该研究提出DRIFT对抗性补丁攻击方法,针对流匹配VLA模型pi0等,通过攻击首个去噪步骤破坏其去噪轨迹,在LIBERO套件上高效破解原本可解决的机器人任务,效果优于相关基线攻击。
中文摘要 AI 辅助
流匹配视觉-语言-动作(VLA)模型(如pi0)通过整合学习到的去噪速度场生成机器人动作,据报道这类模型能抵御易欺骗自回归VLAs的对抗性扰动。我们证明这种鲁棒性在很大程度上是虚幻的:它源于之前的攻击忽略了多步去噪常微分方程(ODE)。我们提出DRIFT(Denoising Redirection via Input perturbation of the Flow-matching Trajectory,即通过流匹配轨迹的输入扰动实现去噪重定向),这是一种放置在机器人夹具上的测试时通用对抗性补丁,用于攻击现成策略的去噪速度场。我们的核心发现有悖直觉:仅攻击第一个去噪步骤比攻击更宽的步骤窗口效果更强且成本更低,我们通过输入空间优化特有的梯度冲突解释这一现象,且该现象与训练时的后门机制完全相反。在四个LIBERO套件上针对pi0和pi0.5,DRIFT用单个小补丁破坏了几乎所有原本可解决的任务,效果远超动作空间和嵌入空间的攻击基线。
英文摘要
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. We introduce DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy. Our central finding is counterintuitive: attacking only the first denoising step is both stronger and cheaper than attacking a wider window of steps, which we explain through a gradient conflict unique to input-space optimization and which is exactly opposite to the training-time backdoor regime. On pi0 and pi0.5 across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks with a small single patch, far exceeding action- and embedding-space attack baselines.