发表机构
University of Western Ontario(西安大略大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现视觉语言模型在极小扰动($\epsilon\leq 4/255$)下并非鲁棒,通过白盒表示对齐实现目标语义替换,图像和视频上分别达到38%和35.9%的完全替换率,并揭示LLM的语义融合现象。
AI 中文摘要
视觉语言模型(VLMs)被广泛部署于安全关键场景中,理解它们能在多大程度上被对抗性扰动所控制,是评估其可信度的前提。现有的表示对齐攻击(representation-alignment attacks)使VLM感知目标图像,在 $\varepsilon \leq 4/255$ 时成功率有限。因此,VLMs在此范围内似乎对扰动具有鲁棒性。我们表明这种鲁棒性并不成立,因为目标语义替换(targeted semantic substitution)在同一范围内能够成功。具体而言,我们将源图像的每个流与目标图像在受害VLM的合并后token空间中的对应流对齐,在白盒威胁模型下进行操作。我们在严格的成功标准下进行评估,要求模型同时说出目标、确认其存在并否认源。在图像中,目标语义在 $\varepsilon = 2/255$ 时出现,完全替换在 $\varepsilon = 4/255$ 时达到38%。在视频中,完全替换在 $\varepsilon = 1/255$ 时达到35.9%。我们还观察到一种“语义融合”(semantic fusion)现象,即大语言模型(LLM)将矛盾的视觉信号合理化为一贯的叙述。
英文摘要
Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial perturbation is a prerequisite for evaluating their trustworthiness. Existing representation-alignment attacks, which make a VLM perceive a target image, achieve limited success at $\varepsilon \leq 4/255$. Therefore, VLMs seems robust to perturbations in this range. We show that this robustness does not hold, as targeted semantic substitution succeeds within the same range. Specifically, we align each stream of the source image with its counterpart in the target image in the victim VLM's post-merger token space, operating under a white-box threat model. We evaluate under a strict success criterion, requiring the model to simultaneously name the target, confirm its presence, and deny the source. In images, target semantics appear at $\varepsilon = 2/255$ and complete replacement reaches 38% at $\varepsilon = 4/255$. On video, complete replacement reaches 35.9% at $\varepsilon = 1/255$. We also observe a phenomenon of semantic fusion, where Large Language Model (LLM) rationalizes contradictory visual signals into a coherent narrative.