arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03498cs.ROcs.AI

检测与抑制:针对VLA模型中对抗性补丁的机制性防御

Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models

  • Keio University(庆应义塾大学)
  • Waseda University(早稻田大学)
  • The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki, Bum Jun Kim, Taiki Miyanishi

AI总结:

本文通过稀疏自编码器识别VLA模型中与对抗性补丁相关的内部特征,并在检测到攻击时条件性抑制该特征,从而在LIBERO-10上提升鲁棒性且不损害正常性能。

AI中文摘要:

对抗性补丁可以通过操纵视觉观察来干扰视觉-语言-动作(VLA)模型,导致机器人控制失败。然而,这些失败背后的内部机制以及有针对性的干预如何缓解它们仍然知之甚少。在本工作中,我们使用稀疏自编码器(SAE)对VLA表示进行机制性分析,并识别出一个特征,其激活与对抗性补丁的存在强相关。基于这一分析,我们仅在线性探针检测到攻击时,在推理时抑制该识别出的特征。这种干预在不增加VLA微调成本的情况下提高了鲁棒性。我们在LIBERO-10上针对VLA对抗性补丁攻击评估了我们的方法。条件性干预在间歇性攻击下提高了成功率,而持续应用相同的干预则显著降低了策略性能。这些结果表明,与攻击相关的内部表示可以为VLA对抗性防御提供有用的目标,并且控制干预时机对于限制对名义策略行为的干扰至关重要。

英文摘要:

Adversarial patches can disrupt Vision-Language-Action (VLA) models by manipulating visual observations, leading to failures in robot control. However, it remains poorly understood which internal mechanisms underlie these failures and how targeted interventions can mitigate them. In this work, we mechanistically analyze VLA representations using a sparse autoencoder (SAE) and identify a feature whose activation strongly correlates with the presence of an adversarial patch. Based on this analysis, we suppress the identified feature at inference time only when a linear probe detects an attack. This intervention improves robustness without the cost of fine-tuning the VLA. We evaluate our method against VLA adversarial patch attacks on LIBERO-10. Conditional intervention improves success rate under intermittent attacks, whereas continuously applying the same intervention substantially degrades policy performance. These results show that attack-related internal representations can provide useful targets for VLA adversarial defense and that controlling when to intervene is important for limiting disruption to nominal policy behavior.

↑