量化实现针对恶意服务提供商的私有稠密检索
Quantization Enables Private Dense Retrieval against Malicious Service Providers
浏览论文内容
中文总结 AI 辅助
针对恶意服务器,提出两轮密码协议实现私有稠密检索,利用低比特量化降低计算成本,实验表明三比特量化在分钟级延迟下保持检索质量。
中文摘要 AI 辅助
稠密检索是检索增强生成(RAG)的关键组件,通过比较查询与大规模语料库中段落的稠密向量表示来检索最相关的文档。在隐私敏感的应用中,服务器会观察查询并控制返回的证据,从而带来机密性和完整性风险。我们将私有稠密检索定义为针对恶意服务器提供查询隐私和检索完整性,并开发了一个两轮密码学协议,该协议同时提供这两种保证。我们的协议将私有且可验证的检索归结为将承诺矩阵与加密向量相乘,并利用低比特量化使该计算变得实用。我们评估了在六种嵌入模型、四种语言模型以及多达268万段落的语料库上,密码学成本、检索质量和下游RAG准确性之间的权衡。结果表明,使用裁剪量化器时,三比特量化在很大程度上保留了检索质量和下游准确性,而针对临床参考规模语料库的私有查询需要一到三分钟的服务器时间。这些结果表明,当分钟级延迟可接受时,私有稠密检索对于中等规模、隐私敏感的语料库已经实用。
英文摘要
Dense retrieval, the key component of Retrieval Augmented Generation (RAG), retrieves the most relevant documents by comparing dense vector representations of queries and passages from a large corpus. In privacy-sensitive applications, the server observes the query and controls which evidence is returned, creating both confidentiality and integrity risks. We formulate private dense retrieval as providing query privacy and retrieval integrity against a malicious server, and develop a two-round cryptographic protocol that provides both guarantees. Our protocol reduces private and verifiable retrieval to multiplication of a committed matrix by an encrypted vector and uses low-bit quantization to make this computation practical. We evaluate the resulting trade-off between cryptographic cost, retrieval quality, and downstream RAG accuracy across six embedding models, four language models, and corpora of up to 2.68 million passages. Our results show that, with a clipped quantizer, three-bit quantization largely preserves retrieval quality and downstream accuracy, while a private query over a corpus the size of a clinical reference requires one to three minutes of server time. These results suggest that private dense retrieval is already practical for moderately sized, privacy-sensitive corpora when minute-scale latency is acceptable.
发表机构
- École de technologie supérieure(高等技术学院)
- Mila(米拉研究所)
- Eurecom(欧洲通信学院)
- Université du Québec à Montréal(魁北克大学蒙特利尔分校)
机构由 AI 辅助整理,请以论文原文为准。