arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于成功与简洁:对可转移视觉语言攻击管道的再审视

On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline

Yuchen Ren, Zhengyu Zhao, Chenhao Lin, Bo Yang, Chao Shen

arXiv 2607.14974首次发表:更新:

发表机构

School of Cyber Science and Engineering, Xi’an Jiaotong University; State Key Laboratory of Mathematical Engineering and Advanced Computing, Information Engineering University(西安交通大学网络空间安全学院; 信息工程大学数学工程与先进计算国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言预训练模型对抗攻击,指出现有复杂攻击管道可简化。提出SimVLA管道解决问题,经实验验证在可转移性和效率上优于基线,强调利用领域知识的重要性,为未来扩展提供简单有效主干。

AI 中文摘要

视觉语言预训练模型(VLPMs)容易受到对抗攻击。近期对VLPMs的可转移攻击遵循复杂损失函数或多阶段文本/图像攻击的通用管道。本文表明复杂攻击管道可更简单且更成功。识别出由不当跨模态交互和过多操作导致的三个被忽视问题,提出简单视觉语言攻击(SimVLA)管道,提高了可转移性和效率。在四个数据集和三个下游任务上实验验证了其优越性,如在Flickr30k文本图像检索数据集上,SimVLA在R@1可转移性上比SOTA基线高出8.01%-14.71%,同时仅消耗约35.73%的时间和46.26%的最大VRAM。突出了利用领域知识的重要性,盲目追求复杂操作可能有害,希望SimVLA能成为未来扩展的简单有效主干。代码可通过链接获取。

英文摘要

Vision-Language Pre-training Models (VLPMs) are known to be vulnerable to adversarial attacks. Recent transferable attacks on VLPMs have followed a common pipeline with complicated loss functions or multi-stage text/image attacks. However, in this paper, we demonstrate that such a sophisticated attack pipeline can be simpler yet more successful. Specifically, we identify three previously overlooked issues caused by inappropriate cross-modal interactions and excessive operations. To address them, we propose the Simple Vision-Language Attack (SimVLA) pipeline, which observably improves transferability and efficiency. Experiments on four datasets and three downstream tasks validate the superiority of our pipeline. For instance, on Flickr30k text-image retrieval dataset, our SimVLA outperforms the SOTA baseline in R@1 transferability by 8.01\%-14.71\%, while consuming only about 35.73\% of the time and 46.26\% of the max VRAM. Overall, the superiority of our SimVLA highlights the importance of leveraging domain knowledge (e.g., our proposed cross-modal word identification), while blindly pursuing intricate operations (e.g, complex loss functions and redundant multi-stage designs) may even be harmful. We hope our SimVLA can serve as a simple yet effective backbone for future extensions. Code is available at https://github.com/RYC-98/SimVLA.

CommentsAccepted for publication in IEEE Transactions on Information Forensics and Security (TIFS)

DOI:10.1109/TIFS.2026.3714129

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑