面向自动驾驶黑盒视觉语言模型的可迁移时空一致性对抗攻击
Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving
浏览论文内容
中文总结 AI 辅助
提出面向自动驾驶黑盒视觉语言模型的时空一致性对抗攻击,通过三阶段方法生成可迁移扰动,实验表明现有视频语言模型高度易受攻击,亟需鲁棒防御。
中文摘要 AI 辅助
视觉语言模型(VLMs)在敏感系统中的快速集成引入了关键的安全漏洞,而现有研究尚未对此进行探索。虽然针对基于图像的模型的对抗攻击鲁棒性已被广泛研究,但在驾驶情境下,VLMs对针对视频的时间感知对抗攻击的敏感性构成了一种独特且未被充分研究的威胁。在本文中,我们提出了一种针对自动驾驶场景中使用的VLM模型的新型视频对抗攻击,名为时空一致性对抗攻击(STCA)。我们的攻击包含三个阶段:模态扩展、空间攻击和STCA攻击。在模态扩展中,我们提出了一种基于字幕引导的帧选择方法,以确保对抗扰动针对最具语义重要性的帧。在空间攻击中,我们制作有效的扰动并保持高相似性。然后,将生成的扰动视频输入STCA阶段,该阶段使用运动引导掩码破坏跨帧的时间一致性。我们的方法在黑盒威胁模型下针对目标VLM模型运行,仅依赖于从白盒替代模型的迁移性。我们在BDD100K和nuScenes自动驾驶数据集上,对三个VLM模型进行了实验:Video LLaVA-7B、Qwen2.5-VL-7B和Dolphin。实验结果表明,空间攻击在实现高SSIM的同时达到了较高的攻击成功率(ASR)。我们的发现揭示,现有的视频语言模型在自动驾驶场景中仍然高度易受对抗攻击的影响,这强调了为VLM模型建立鲁棒防御的紧迫性。
英文摘要
The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in exist studies. While adversarial attack robustness has been extensively studied for image-based models, the susceptibility of VLMs to temporally-aware adversarial attacks against video in driving context poses a distinct and under examined threat. In this paper, we introduce novel adversarial attack against video targeting VLM models used for autonomous driving scenes named Spatial Temporal Coherence Adversarial Attack (STCA). Our attack comprise from three stages: modalities expansion, Spatial attack, and STCA attack. In modalities expansion, we propose caption-guided frame selection method in order to ensure that adversarial perturbation target the most semantically significant frames. Secondly.In spatial attack, we craft effective perturbation and preserve high similarity. Then the perturbed video generated fed into STCA stage that disrupt cross-frame temporal coherence using motion guided mask. Our method operate under black box threat model against victim target VLMs, relying solely on transferability from white-box surrogate model.We conduct our experiments on the BDD100K and nuScenes autonomous driving datasets across three VLM models: Video LLaVA-7B, Qwen2.5-VL-7B, and Dolphin. Experimental results demonstrate spatial attack achieves an ASR with high SSIM. Our finding reveal that existing video language model, remain highly susceptible to adversarial attack in autonomous driving scenarios, underscoring the urgent need for robust defense for VLM models.