arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FreqDoor:面向视觉-语言模型后门攻击的频率域隐藏木马

FreqDoor: A Hidden Trojan in the Frequency Domain for Backdoor Attacks on Vision-Language Models

Yasir Arafat Prodhan, Sadad Hasan, Mohammed Imamul Hassan Bhuiyan

arXiv 2609.07048首次发表:更新:

AI 中文总结

提出FreqDoor,一种在频率域植入不可见触发器的训练时后门攻击,通过混合幅度谱并保留相位,在多个VLM上实现高攻击成功率。

AI 中文摘要

视觉-语言模型(VLM)近期在开放式图像到文本生成方面展现出卓越进展。然而,其多模态特性使其持续易受后门攻击。现有的VLM后门触发器要么是空间域的、文本域的,要么是双模态的,这可能导致局部化或可识别的触发器模式。在本工作中,我们探索了不同的攻击面,并提出了FreqDoor,一种在频率域植入触发器的训练时后门攻击。FreqDoor选择性地混合来自触发器源图像的幅度谱分量,同时保留干净图像的相位,以生成空间分布且视觉上不可感知的触发器,而无需修改文本输入。我们在BLIP-2、InstructBLIP和LLaVA上评估了该攻击在图像描述和视觉问答任务中的表现。在Flickr8k上,FreqDoor在三个模型上分别达到了99.6%、99.8%和98.4%的攻击成功率,同时保持了生成描述语义质量。在VQAv2上,相应的攻击成功率分别为99.6%、92.4%和79.6%。

英文摘要

Vision-language models (VLMs) have recently shown excellent progress in open-ended image-to-text generation. However, their multimodal nature makes them persistently vulnerable to backdoor attacks. Existing backdoor triggers for VLMs are either spatial, textual, or bimodal, which may yield localized or recognizable trigger patterns. In this work, we explore a different attack surface and propose \ textsc {FreqDoor}, a training-time backdoor attack that implants triggers in the frequency domain. \ textsc {FreqDoor} mixes amplitude-spectrum components from a trigger-source image selectively while preserving the phase of a clean image to generate a spatially distributed and visually imperceptible trigger without modifying the textual input. We evaluate the attack on BLIP-2, InstructBLIP, and LLaVA for image captioning and visual question answering. On Flickr8k, \ textsc {FreqDoor} achieves attack success rates of $99.6\%$, $99.8\%$, and $98.4\%$ on the three models, respectively, while preserving the semantic quality of the generated captions. On VQAv2, the corresponding attack success rates are $99.6\%$, $92.4\%$, and $79.6\%$.

Comments13 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑