arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DSPrompt:针对M-RAG污染的动态软提示防御

DSPrompt: Dynamic Soft Prompt Defense Against M-RAG Corruption

Chang Liu, Yuni Lai, Mingyue Cui, Cong Tian, Yunyan Zhang, Xian Wu, Kai Zhou, Bin Xiao

arXiv 2608.16536首次发表:更新:

发表机构

Xidian University; The Hong Kong Polytechnic University; Tencent Jarvis Lab(西安电子科技大学; 香港理工大学; 腾讯Jarvis Lab)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对M-RAG面临的对抗攻击问题,提出DSPrompt动态软提示防御框架,通过在冻结检索器编码器插入可学习软提示,以低开销实现了优于现有方法的防御效果,同时维持检索与生成性能。

AI 中文摘要

多模态检索增强生成(Multimodal Retrieval Augmented Generation,M-RAG)日益受到对抗攻击的威胁,攻击者会构造恶意数据以生成与向量空间中良性条目对齐的嵌入,从而欺骗检索环节并诱导模型输出有害内容。现有防御方法主要在查询阶段运作,依赖辅助检测器、相似度重排序或特征一致性检查,但这些方法存在显著的推理开销、对未见过的攻击策略泛化能力差,且通常假设特定的攻击分布。为解决上述问题,本文提出DSPrompt,一种动态软提示防御框架,该框架无需修改检索流程,直接重塑检索器的嵌入语义。它在冻结检索器的视觉和文本编码器的每一层插入少量可学习的软提示,采用从浅层到深层的长度调度,以适应模型层的容量。这些提示在动态最小-最大方案下进行训练:在线多模态攻击者针对当前检索器持续构造硬对抗文档,而防御方则进行更新,以将此类文档排除在top-k之外,同时保持良性证据的排序和多样性。由于防御后的编码器可像标准密集检索一样预先计算并建立索引,DSPrompt不会产生额外的每查询优化开销,且引入的参数少于1%。在四个基准数据集和三种代表性投毒攻击上开展的大量实验表明,DSPrompt大幅降低了攻击成功率和投毒检索率,同时保持接近无损的检索效用和生成保真度,以远低于现有防御基线的计算成本持续超越它们。

英文摘要

Multimodal Retrieval Augmented Generation (M-RAG) is increasingly vulnerable to adversarial attacks where malicious data are crafted to produce embeddings that align with benign entries in the vector space, deceiving retrieval and inducing harmful outputs. Existing defenses primarily operate at query time, relying on auxiliary detectors, similarity re-ranking, or feature-consistency checks. However, these approaches suffer from non-trivial inference overhead, generalize poorly to unseen attack strategies, and often assume specific attack distributions. To address this, we propose DSPrompt, a Dynamic Soft Prompt defense framework that directly reshapes the retriever's embedding semantics, without modifying the retrieval pipeline. It inserts few learnable soft prompts into each layer of the visual and textual encoders of a frozen retriever, utilizing a shallow-to-deep length schedule that is adaptive to the capacity in the model layers. These prompts are trained under a dynamic min-max scheme: an online multimodal attacker continually crafts hard adversarial documents against the current retriever, while the defender is updated to push such documents out of the top-k while preserving the ranking and diversity of benign evidence. Because the defended encoder can be pre-computed and indexed exactly as in standard dense retrieval, DSPrompt incurs no additional per-query optimization and introduces fewer than 1% additional parameters. Extensive experiments across four benchmarks and three representative poisoning attacks show that DSPrompt substantially reduces the attack success rate and poison retrieval rate while maintaining near-lossless retrieval utility and generation fidelity, consistently outperforming existing defense baselines at a fraction of their computational cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑