arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06977cs.CVcs.AI

视觉不变性增强的特征最优对齐:针对闭源多模态大语言模型的可迁移对抗攻击

Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs

Xiaojun Jia, Simeng Qin, Yiming Li, Jie Liao, Sensen Gao, Ke Ma, Yang Liu, Xiaochun Cao

首次发表
浏览论文内容

中文总结 AI 辅助

针对闭源多模态大语言模型的黑盒攻击,提出视觉不变性增强的特征最优对齐方法(IAU-FOA),通过全局与局部对齐及自适应非平衡传输,提升定向攻击的可迁移性,实验验证其优于现有方法。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)仍然容易受到可迁移对抗样本的攻击,尤其是在仅能访问开源替代模型的黑盒设置中。现有的定向迁移攻击主要利用全局图像级特征(如编码器的[CLS]嵌入)来对齐对抗样本和目标样本。然而,这种粗粒度的对齐未能充分利用补丁级的视觉结构,限制了其在异构闭源MLLMs之间的可迁移性。我们提出了IAU-FOA,一种结合自适应非平衡传输的视觉不变性增强特征最优对齐攻击,以提高对闭源MLLMs的定向可迁移性。IAU-FOA在全局和局部两个层面同时对齐对抗样本和目标样本:基于余弦的目标函数缩小了它们的全局语义差距,而补丁标记被聚类为紧凑的局部模式,并通过最优传输进行细粒度特征对齐。平衡最优传输即使对于缺乏可靠对应关系的局部聚类也强制执行固定的边缘质量,这可能会引入误导性的对齐梯度。因此,我们引入了置信度自适应的非平衡传输,以放宽对弱匹配聚类的这些约束,旨在减少不可靠的局部对齐并提高对抗可迁移性。我们进一步研究了输入变换的影响,并提出了视觉不变性增强,它应用双向像素强度缩放和逐通道白平衡调整来模拟曝光、对比度、光照和色温的变化。该策略鼓励对抗扰动在不同视觉编码器之间泛化。在开源和闭源MLLMs上进行的大量实验表明,IAU-FOA始终优于最先进的可迁移攻击方法。代码可在https URL获取。

英文摘要

Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-source surrogate models are accessible. Existing targeted transfer attacks mainly align adversarial and target samples using global image-level features, such as encoder [CLS] embeddings. However, such coarse alignment insufficiently exploits patch-level visual structures, limiting transferability across heterogeneous closed-source MLLMs. We propose IAU-FOA, a visual-invariance-augmented feature optimal alignment attack with adaptive unbalanced transport, to improve targeted transferability against closed-source MLLMs. IAU-FOA aligns adversarial and target samples at both global and local levels: a cosine-based objective narrows their global semantic gap, while patch tokens are clustered into compact local patterns and matched through optimal transport for fine-grained feature alignment. Balanced optimal transport enforces fixed marginal masses even for local clusters without reliable counterparts, potentially introducing misleading alignment gradients. We therefore introduce confidence-adaptive unbalanced transport to relax these constraints for weakly matched clusters, aiming to reduce unreliable local alignment and improve adversarial transferability. We further study the effect of input transformations and propose visual-invariance augmentation, which applies bidirectional pixel-intensity rescaling and per-channel white-balance adjustment to simulate exposure, contrast, illumination, and color-temperature variations. This strategy encourages adversarial perturbations to generalize across different visual encoders. Extensive experiments on open-source and closed-source MLLMs show that IAU-FOA consistently outperforms state-of-the-art transferable attack methods. Code is available at https://github.com/jiaxiaojunQAQ/IAU-FOA.

发表机构

  • Nanyang Technological University(南洋理工大学)
  • Northeastern University(东北大学)
  • Chongqing University(重庆大学)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

↑