禁止你的注意力:通过在频域中有选择地移除固有焦点来欺骗多模态大语言模型
Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain
浏览论文内容
中文总结 AI 辅助
该研究针对多模态大语言模型,提出一种频域相位感知的对抗攻击方法,结合对抗提示学习,可有效欺骗模型,效果优于现有攻击。
中文摘要 AI 辅助
多模态大语言模型(Multimodal Large Language Models,MLLMs)扩展了大语言模型(Large Language Models,LLMs)处理更多上下文多模态信息的能力,在各类现实多模态应用中展现出显著进展。尽管MLLMs具备强大的感知与推理能力,近期研究表明它们仍极易受到对抗性输入的影响,尤其是针对视觉组件的对抗性输入。然而,现有攻击方法主要聚焦于全局扰动,缺乏对MLLMs内部如何解释视觉结构的理解。本文尝试探究MLLMs在频域中的固有焦点,发现其预测对相位信息尤为敏感,而相位信息编码了关键的结构与语义线索。基于这一观察,我们提出了一种新颖的感知相位的对抗攻击框架,该框架明确将对抗性扰动限制在与结构相关的相位区域,以抑制MLLMs的焦点,实现有效且不易察觉的攻击。为进一步放大结构影响,我们还引入了一个辅助的对抗性提示学习模块,以引导相位敏感区域周围的多模态错位,误导MLLMs的注意力指向目标结构模式。在多个代表性MLLM模型和数据集上开展的大量实验表明,与现有攻击方法相比,我们的方法具有更优的有效性。
英文摘要
Multimodal large language models (MLLMs) have extended the capability of large language models (LLMs) to process more contextual multimodal information, showing remarkable progress in diverse realistic multimodal applications. Despite their strong perception and reasoning abilities, recent studies reveal that MLLMs remain highly vulnerable to adversarial inputs, especially those targeting visual components. However, existing attacks mainly focus on global perturbations, lacking an understanding of how MLLMs internally interpret visual structures. In this paper, we make the attempt to investigate the intrinsic focus of MLLMs in the frequency domain and discover that their predictions are particularly sensitive to phase information, which encodes essential structural and semantic cues. Based on this observation, we propose a novel phase-aware adversarial attack framework that explicitly restricts adversarial perturbations to structure-relevant phase regions to suppress the MLLMs' focus for effective and imperceptible attacks. To further amplify the structural influence, we also introduce an auxiliary adversarial prompt learning module to guide multimodal misalignment around phase-sensitive regions, misleading the MLLM's attention toward targeted structural patterns. Extensive experiments on multiple representative MLLM models and datasets demonstrate the superior effectiveness of our method compared to existing attacks.
发表机构
- Institute for Math & AI, Wuhan University(武汉大学数学与人工智能学院)
- College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
- School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)
- Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
- Cyberspace Institute of Advanced Technology, Guangzhou University(广州大学先进技术网络空间学院)
- College of Computer and Information Engineering, Zhejiang Gongshang University(浙江工商大学计算机与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。