arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33833cs.CVcs.LG

一击愚弄所有模型:针对前沿多模态大语言模型的高迁移性黑盒对抗攻击

One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs

Sen Nie, Jie Zhang, Zhongqi Wang, Shiguang Shan, Xilin Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对前沿多模态大语言模型,提出O-Attack黑盒攻击框架,利用替代模型的跨模态语义空间,通过语义共识优化扰动,显著提升攻击成功率与迁移性,并揭示实际安全风险。

中文摘要 AI 辅助

对抗攻击长期以来一直是机器学习系统面临的根本性威胁。随着多模态大语言模型(MLLMs)的快速演进和广泛部署,评估其对此类攻击的脆弱性对于其安全使用至关重要。在本工作中,我们研究单个对抗性图像是否能在黑盒设置下持续误导多样化的前沿MLLMs。我们提出了O-Attack,一种高迁移性的黑盒攻击框架。该框架基于我们的洞察:替代模型包含一个广泛的、高层次的、跨模态对齐的语义空间。这个空间超越了最终层输出,提供了多个语义一致的表征,这些表征尚未被现有攻击充分利用。在此空间内,O-Attack锚定对齐表征,逐步拓宽语义条件,并通过语义共识优化扰动以促进一致的目标对齐。通过利用与M-Attack相同的替代模型充分挖掘此空间,O-Attack在GPT-5.4(从29.1%提升至77.2%)、Claude-4.6(从42.8%提升至81.6%)和Gemini-3.1(从38.2%提升至80.9%)上的攻击成功率得到提升。在24个MLLMs上的广泛实验表明,O-Attack在黑盒迁移性方面优于六种最先进的方法,在提示词上具有一致的有效性,并提高了效率和不可感知性。本工作揭示了黑盒对抗攻击对前沿MLLMs构成的实际安全风险,强调了进行更严格的鲁棒性评估和更有效防御的必要性。

英文摘要

Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. We propose O-Attack, a highly transferable black-box attack framework. This framework builds on our insight that surrogate models contain a broad, high-level, cross-modally aligned semantic space. This space extends beyond final-layer outputs and provides multiple semantically consistent representations that remain underexploited by existing attacks. Within this space, O-Attack anchors aligned representations, progressively broadens semantic conditions, and optimizes perturbations through semantic consensus to promote consistent target alignment. By fully exploiting this space with the same surrogate models as M-Attack, O-Attack raises attack success rates on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Extensive experiments across 24 MLLMs show that O-Attack outperforms six state-of-the-art methods in black-box transferability, with consistent effectiveness across prompts and improved efficiency and imperceptibility. This work exposes the practical safety risks posed by black-box adversarial attacks against frontier MLLMs, underscoring the need for more rigorous robustness evaluation and more effective defenses.

发表机构

  • State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所人工智能安全国家重点实验室)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑