发表机构
The University of Queensland; CSIRO’s Data 61; City University of Hong Kong(昆士兰大学; 联邦科学与工业研究组织数据61; 香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对向量数据库中文档嵌入易受嵌入反演攻击的问题,提出SHAQ防御方法,通过生成影子查询编码后存储,在保留检索效用的同时提升隐私。
AI 中文摘要
大型语言模型(LLMs)越来越依赖信息检索(IR)系统,如检索增强生成(RAG),以融入特定领域知识而无需代价高昂的重新训练。这些系统通常将预计算的文档嵌入存储在基于云的向量数据库中。然而,此类嵌入易受嵌入反演攻击(EIAs)的影响,此类攻击可重建其底层文本。现有防御措施,如添加噪声或缩放嵌入,往往提供有限的隐私保护或显著降低检索效用。我们提出SHAQ(影子查询生成),一种针对EIAs的语义分解与嵌入解耦防御方法。SHAQ基于EIAs依赖嵌入与其原始文本间强耦合的洞察,并非直接存储文档嵌入,而是使用生成式语言模型创建多样的影子查询,捕获每个文档的不同语义方面。这些查询被编码并存储以替代原始文档嵌入,从而分解文档语义并将存储的嵌入与源文本解耦。在多样IR数据集上的实验表明,SHAQ在保留检索效用的同时大幅提升隐私,实现低至0.2104的恢复率,比基线防御多防御高达19.50%的token,达到高达0.7967的MAP@10并实现高达5.53%的效用提升。这些结果表明,语义分解与嵌入解耦为防御EIAs提供了直接修改嵌入的有效替代方案。
英文摘要
Large language models (LLMs) increasingly rely on information retrieval (IR) systems, such as Retrieval-Augmented Generation (RAG), to incorporate domain-specific knowledge without costly re-training. These systems often store pre-computed document embeddings in cloud-based vector databases. However, such embeddings are vulnerable to embedding inversion attacks (EIAs), which can reconstruct their underlying text. Existing defenses, such as adding noise or scaling embeddings, often provide limited privacy or significantly reduce retrieval utility. We propose SHAQ (shadow query generation), a semantic-decomposition and embedding-decoupling defense against EIAs. SHAQ is based on the insight that EIAs rely on the strong coupling between an embedding and its original text. Instead of storing document embeddings directly, SHAQ uses a generative language model to create diverse shadow queries that capture different semantic aspects of each document. These queries are then encoded and stored in place of the original document embeddings, thereby decomposing document semantics and decoupling stored embeddings from the source text. Experiments across diverse IR datasets show that SHAQ substantially improves privacy while preserving retrieval utility, achieving a recovery rate as low as 0.2104, defending up to 19.50% more tokens than baseline defenses, and reaching up to 0.7967 MAP@10 with up to 5.53% utility improvement. These results demonstrate that semantic decomposition and embedding decoupling provide an effective alternative to directly modifying embeddings for defending against EIAs.
Comments13 pages