AdvNav:视觉语言导航中基于行为引导的黑盒对抗攻击
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究视觉语言导航系统对抗攻击,提出AdvNav框架,利用双粒度行为反馈构建替代目标,采用混合优化策略。在R2R数据集上评估,对两种模型取得较高攻击成功率,证明其有效性与通用性,揭示感知漏洞并为VLN模型设计提供见解。
AI中文摘要:
尽管具身人工智能取得了进展,但视觉语言导航系统仍易受对抗性视觉干扰影响。大多数现有方法依赖白盒访问目标模型梯度,对实际部署系统不现实且计算量大。以往黑盒方法主要针对单步瞬时决策任务,难以处理任务复杂性和时间依赖性。因此提出AdvNav,一种行为引导的黑盒对抗攻击框架,在导航中干扰智能体的第一人称视图。设计了双粒度基于行为的反馈来构建替代目标,指导混合优化策略,通过自适应更新启发式调整扰动强度,通过遗传算法进化噪声空间结构。在R2R数据集上对基于Transformer的HAMT和基于LLM的MapGPT进行评估,AdvNav实现了49.70/65.96/87.30%的攻击成功率。结果证明了AdvNav的有效性和通用性,揭示了关键感知漏洞,并为未来弹性VLN模型的设计提供了见解。
英文摘要:
Despite progress in Embodied AI, Vision-and-Language Navigation systems remain vulnerable to adversarial visual disturbances. Most existing methods rely on white-box access to target model gradients, which is often unrealistic for real-world deployed systems and computationally exhaustive due to recursive backpropagation for optimization, limiting their applicability. While previous black-box methods predominantly target single-step, instantaneous decision tasks, they struggle to handle the task complexities and temporal dependencies. This highlights the need for a gradient-free attack method that can effectively disrupt the multistep sequential perception-action loop using only observable inputs and outputs. Therefore, we propose AdvNav, a behavior-guided black-box adversarial attack framework that disturbs an agent's first-person views during navigation. To construct an informative surrogate objective for effective optimization guidance in gradient-free search under the black-box setting, we design a dual-granularity behavior-based feedback, aggregating a trajectory-level performance score representing overall navigation degradation, an action-level reward score considering the potential decision risk, and a deviation indicator, all of which are extracted from the agent's self-output behaviors. This feedback guides a hybrid optimization strategy that heuristically tunes perturbation strength via adaptive updates and evolves noise spatial structure genetically, to iteratively discover the most disruptive noise configuration. Evaluated against Transformer-based HAMT and LLM-based MapGPT with two types of backbones on R2R dataset, AdvNav achieves 49.70/65.96/87.30% Attack Success Rate. The result demonstrates the effectiveness and generality of AdvNav, reveals critical perception vulnerabilities and offers insights for the design of future resilient VLN models.