arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18938cs.AIcs.LG

打破最弱环节以规避视觉语言模型

Breaking the weakest link to evade vision language models

Ilan Zini, Boussad Addad, Katarzyna Kapusta

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对视觉语言模型(VLMs)提出仅优化视觉编码器的梯度攻击方法,在Qwen2.5-VL等模型上验证微小扰动可规避模型,凸显其对抗操纵的脆弱性并呼吁强化安全机制。

中文摘要 AI 辅助

视觉语言模型(Vision Language Models,VLMs)近期已成为多模态AI系统的关键组成部分,支持在现实场景及安全关键应用中对视觉与文本输入进行联合推理。尽管其部署日益广泛,但VLMs对抗对抗性威胁的鲁棒性仍未得到充分探究,尤其是针对多模态对齐的规避攻击场景。本研究探究了VLMs对应用于视觉输入的对抗性扰动的脆弱性,并研究两类攻击设置:无目标攻击,目标为破坏模型对原始图像的解读;有目标攻击,攻击者旨在迫使模型生成与原始图像无关的特定语义描述。为高效生成对抗样本,我们提出一种基于梯度的攻击方法,该方法仅在VLM的视觉编码器上执行优化,而非在整个多模态架构上进行。此设计显著降低了攻击的计算成本与资源需求,同时保持了较强的有效性。我们在多个开源VLMs上评估了该方法,包括Qwen2.5-VL、Granite-Vision、FastVLM及Phi-3.5-Vision,结果表明微小的、人类难以察觉的扰动可大幅改变模型生成的文本解读。我们的发现凸显了现代VLMs对抗对抗性操纵的脆弱性,强调了多模态AI系统中改进鲁棒性与安全机制的必要性。

英文摘要

Vision Language Models (VLMs) have recently emerged as a critical component of multimodal AI systems, enabling joint reasoning over visual and textual inputs in real-world and safety-critical applications. Despite their growing deployment, the robustness of VLMs against adversarial threats remains insufficiently explored, particularly in the context of evasion attacks targeting multimodal alignment. In this work, we investigate the vulnerability of VLMs to adversarial perturbations applied to visual inputs and study two attack settings: untargeted attacks, where the goal is to disrupt the model's interpretation of the original image, and targeted attacks, where the adversary aims to force the model to generate a specific semantic description unrelated to the original image. To efficiently generate adversarial examples, we propose a gradient-based attack method that performs optimization exclusively on the vision encoder of the VLM rather than on the entire multimodal architecture. This design significantly reduces the computational cost and resource requirements of the attack while maintaining strong effectiveness. We evaluate our approach on several open-source VLMs, including Qwen2.5-VL, Granite-Vision, FastVLM, and Phi-3.5-Vision, and show that small, human-imperceptible perturbations can substantially alter the textual interpretation produced by the models. Our findings highlight the vulnerability of modern VLMs to adversarial manipulation and emphasize the need for improved robustness and security mechanisms in multimodal AI systems.

发表机构

  • ESILV(ESILV高等工程师学院)
  • Thales(泰雷兹集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑