发表机构
Department of Mathematics, Hong Kong Baptist University; School of Automation, Northwestern Polytechnical University(香港浸会大学数学系; 西北工业大学自动化学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大型视觉语言模型在遥感图像解释中的对抗攻击问题,提出GeoThreat方法,通过在概念和感知层面调制对抗表示,结合对抗目标相似性梯度等技术,实现可转移目标对抗攻击,实验证明该方法在可转移性和可控性上具有优越性。
AI 中文摘要
对大型视觉语言模型(LVLMs)的对抗攻击是评估其跨模态语义理解鲁棒性的有效手段。现有研究主要集中在破坏视觉输入以在一般视觉语言任务中诱导预定义的错误响应,而遥感领域的相应研究仍未充分探索。与自然图像理解相比,遥感图像解释需要对局部判别线索和全局场景上下文进行联合推理。这给在黑盒设置下实现对指定响应的可转移语义操纵带来了额外挑战。为应对这些挑战,我们提出了GeoThreat,一种针对LVLMs的用于遥感图像解释的可转移目标对抗攻击方法。具体而言,GeoThreat在概念和感知层面根据目标内容调制对抗表示。代理图像编码器的类令牌用作概念表示,而感知表示则通过协作重要性估计从对抗示例的补丁令牌中提取。除了在各层展开注意力分数外,我们还纳入对抗目标相似性梯度,以更忠实地表征局部视觉线索与预期语义操纵的相关性。然后,感知表示以交叉注意力方式与目标补丁令牌动态对齐,促进局部线索向指定语义细节的适应。最后,通过基于集成的概念校准和感知适应联合优化迭代更新对抗扰动。在各种LVLMs上的大量实验证明了GeoThreat在可转移性和可控性方面的优越性。
英文摘要
Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing their robustness in cross-modal semantic understanding. Existing studies mainly focus on corrupting visual inputs to induce predefined erroneous responses in general vision-language tasks, whereas corresponding investigations in remote sensing fields remain largely underexplored. Compared with natural image understanding, remote sensing image interpretation requires joint reasoning over local discriminative cues and global scene context. This poses additional challenges to achieving transferable semantic manipulation toward specified responses under black-box settings. To tackle these challenges, we propose GeoThreat, a transferable targeted adversarial attack method against LVLMs for remote sensing image interpretation. Specifically, GeoThreat modulates adversarial representations in accordance with the target content at both conceptual and perceptual levels. The class tokens from surrogate image encoders are employed as conceptual representations, while perceptual representations are distilled from patch tokens of the adversarial example through collaborative importance estimation. Beyond merely rolling out attention scores across layers, we incorporate adversarial-target similarity gradients to more faithfully characterize the relevance of local visual cues to the intended semantic manipulation. The perceptual representations are then dynamically aligned with target patch tokens in a cross-attentive manner, facilitating the adaptation of local cues toward designated semantic details. Finally, adversarial perturbations are iteratively updated via ensemble-based joint optimization of conceptual calibration and perceptual adaptation. Extensive experiments across diverse LVLMs demonstrate the superiority of GeoThreat in both transferability and controllability.
CommentsThe code will be released at https://github.com/fuyimin96/GeoThreat upon acceptance