arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

行走嵌入空间:多模态RAG的数据存储提取

Walking the Embedding Space: Datastore Extraction from Multimodal RAG

Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen

arXiv 2610.01871首次发表:更新:

发表机构

University of Groningen; Leiden University(格罗宁根大学; 莱顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出IMMRAG,一种针对图像返回型多模态RAG的黑盒自适应数据提取攻击,通过混合阴影图像与相关性加权重采样,在2500次查询中重建数百张图像,揭示多模态数据安全防护的紧迫性。

AI 中文摘要

多模态检索增强生成(MRAG)已成为一种可靠且经济高效的技术,可将多模态大语言模型(MLLMs)的生成能力锚定到相关、最新且外部知识中。尽管该方法具有诸多优势,如减少幻觉行为,但也引入了新的攻击面,包括隐私信息泄露和易受数据提取攻击的漏洞。在本文中,我们提出了IMMRAG,一种自适应且自动化的数据提取攻击流程,在针对“图像返回型”MRAG的黑盒设置中运行,该配置下检索到的视觉产物本身即为响应。每次查询将攻击者持有的阴影图像与已从系统恢复的图像混合,并通过相关性加权重采样引导后续查询朝向嵌入空间中仍能产生新检索结果的区域。与当前旨在通过将恶意查询作为文本提示来诱导模型泄露数据的提取攻击不同,IMMRAG将恶意指令嵌入用户给定的输入图像中。我们在三个合理且不同的现实场景中评估了IMMRAG:医疗助手、文档聚焦助手和通用工具。实验涉及研究攻击在多个CLIP系列检索器上的有效性,以及不同生成器的影响。单次2500次查询的运行在局部特征对应下可重建多达611张不同的放射学图像、566张文档扫描件和416张通用图像,并且达到非自适应基线5.6倍的不同数据存储项数量。我们的结果表明,迫切需要专门针对多模态数据设计的安全防护措施。

英文摘要

Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce $\immrag$, an adaptive and automatic data extraction attack procedure operating in a black box setting against \emph{image-returning} MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, $\immrag$ embeds the malicious instructions inside a user-given input image. We evaluate $\immrag$ on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to $5.6\times$ as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑