发表机构
The Chinese University of Hong Kong, Shenzhen; State Key Laboratory of Internet Architecture, Tsinghua University(香港中文大学(深圳); 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Spruce将表示与密码协议协同设计,通过紧凑二进制编码、固定半径MPC、私有聚类剪枝等技术,在保持检索质量的同时大幅提升私有外包检索的速度与吞吐量。
AI 中文摘要
检索增强生成(RAG)已使针对大型文档集合的密集检索成为标准构建模块。企业越来越多地将向量索引外包给不可信的云服务,这暴露了专有语料库和用户查询。密码学保护颇具挑战性,因为每次查询都要搜索语料库规模的状态,导致计算、相关随机性和通信量随语料库规模增长。在百万级文档规模下,朴素的安全实现每次查询耗时数分钟,通信量约90 GB;即使是近期优化的系统也需要10-22秒。我们提出Spruce(基于紧凑嵌入的可扩展私有外包检索),它将表示学习与密码协议协同设计。Spruce学习保留全精度重排序候选的紧凑二进制编码,在两方安全多方计算(MPC)下用高效汉明距离计算替代语料库范围的嵌入评分。语料库校准的固定半径协议避免多轮候选选择,同时保持检索质量。Spruce还提供私有聚类剪枝,以轻微质量损失换取计算量大幅减少,以及单核所有者运营的 dealer,消除不经意传输(OT)的云预处理瓶颈。在包含38.3万至542万份文档的四个语料库上,Spruce以仅382至1952的中位候选集保持原始搜索质量;在10 Gbps的服务器间带宽下,全扫描耗时0.21至2.97秒,比最近测量的最接近工作快4.8至6.7倍;私有剪枝耗时0.06至1.09秒,加速比达13.1至22.9倍,保留93.9%至97.3%的全精度NDCG;在最大语料库上,剪枝与dealer共同使每链路1 Gbps的持续吞吐量提升31.5倍。
英文摘要
Retrieval-Augmented Generation (RAG) has made dense retrieval over large document collections a standard building block. Organizations increasingly outsource vector indexes to untrusted clouds, exposing proprietary corpora and user queries. Cryptographic protection is challenging because each query searches corpus-scale state, causing computation, correlated randomness, and communication to grow with the corpus. At million-document scale, a naive secure implementation takes minutes and about 90 GB of communication per query. Even recent optimized systems require 10--22 seconds. We propose Spruce (Scalable Private Outsourced Retrieval Using Compact Embeddings), which co-designs representations with the cryptographic protocol. Spruce learns compact binary codes that preserve candidates for full-precision reranking, replacing corpus-wide embedding scoring with efficient Hamming-distance computation under two-server multi-party computation (MPC). A corpus-calibrated fixed-radius protocol avoids multi-round candidate selection while preserving retrieval quality. Spruce also provides private cluster pruning, which trades minor quality loss for substantially less computation, and a one-core owner-operated dealer that removes cloud OT preprocessing bottlenecks. Across four corpora containing 383K--5.42M documents, Spruce preserves the original search quality with median candidate sets of only 382--1,952. At 10 Gbps inter-server bandwidth, full scans take 0.21--2.97 seconds, $4.8$--$6.7\times$ faster than the closest measured prior work. Private pruning takes 0.06--1.09 seconds, achieves $13.1$--$22.9\times$ speedups, and retains $93.9\%$--$97.3\%$ of full-float NDCG. On the largest corpus, pruning and the dealer jointly improve sustained throughput by $31.5\times$ at 1 Gbps per link.
Comments23 pages, 10 tables, 6 figures