arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14685cs.CR

ViTeGate:面向视觉-语言检索增强生成的视觉-文本触发知识投毒

ViTeGate: Visual-Textual Triggered Knowledge Poisoning for Vision-Language Retrieval-Augmented Generation

Xue Tan, Xuandi Zeng, Yu Shao, Zhongli Fang, Mingyu Luo, Xiaoyan Sun, Ping Chen, Jun Dai

首次发表
浏览论文内容

中文总结 AI 辅助

ViTeGate 提出视觉-文本双触发知识投毒攻击,通过条件性提升投毒证据并诱导指定响应,实现选择性激活,在 InfoSeek 上攻击成功率高达 0.98,同时保持干净答案准确率 0.93。

中文摘要 AI 辅助

现代视觉-语言检索增强生成(VLRAG)系统通过检索到的视觉和文本证据增强大型视觉-语言模型(LVLMs),使其能够基于外部知识生成响应。然而,检索流程也带来了攻击面: adversaries 可以向知识库中注入被投毒的图像-文本对,从而影响模型输出。现有的知识投毒攻击通常是持续激活的,使得被投毒的证据一旦被检索到就会影响生成。这种缺乏精确激活控制的问题使得恶意行为难以被限制在预期输入上,降低了攻击的隐蔽性和有效性。在本文中,我们提出了 ViTeGate,一种针对 VLRAG 系统的视觉-文本触发知识投毒攻击。ViTeGate 使用视觉触发器有条件地将被投毒的证据提升到检索结果中,并使用文本触发器从检索到的证据中诱导攻击者指定的响应。通过协调检索和生成,ViTeGate 在视觉触发器不存在时减少投毒暴露,并在文本触发器不存在时保持正常响应。双触发器设计实现了选择性攻击激活,并减少了意外的单触发器激活。在多个查询数据集、检索器和 LVLMs 上的实验验证了 ViTeGate 的有效性。在 InfoSeek 上,ViTeGate 实现了高达 0.98 的攻击成功率,同时保持了高达 0.93 的干净答案准确率。

英文摘要

Modern Vision-Language Retrieval-Augmented Generation (VLRAG) systems augment Large Vision-Language Models (LVLMs) with retrieved visual and textual evidence, enabling responses grounded in external knowledge. However, the retrieval pipeline also creates an attack surface: adversaries can inject poisoned image-text pairs into the knowledge corpus to influence model outputs. Existing knowledge poisoning attacks are typically always-on, allowing poisoned evidence to affect generation whenever it is retrieved. This lack of precise activation control makes it difficult to confine malicious behavior to intended inputs, reducing both attack stealth and effectiveness. In this paper, we propose ViTeGate, a visual-textual triggered knowledge poisoning attack for VLRAG systems. ViTeGate uses a visual trigger to conditionally promote poisoned evidence into retrieval results and a textual trigger to induce an attacker-specified response from the retrieved evidence. By coordinating retrieval and generation, ViTeGate reduces poison exposure when the visual trigger is absent and preserves normal responses when the textual trigger is absent. The two-trigger design enables selective attack activation and reduces unintended single-trigger activation. Experiments across multiple query datasets, retrievers, and LVLMs validate the effectiveness of ViTeGate. On InfoSeek, ViTeGate achieves an attack success rate of up to 0.98 while maintaining a clean answer accuracy of up to 0.93.

↑