AI 中文总结
该研究提出单GPU上基于全同态加密的十亿级最近邻搜索系统,结合降维与分层结构,在多个数据集上实现高效查询,同时平衡了访问模式泄露与计算开销。
AI 中文摘要
我们构建了一个系统,用于回答“哪些数据库向量与我的查询最相似?”,而服务器全程无法看到查询内容。查询采用全同态加密(FHE);服务器对密文执行所有评分操作,并返回仅客户端可读取的加密结果。面临的挑战是速度:在十亿级向量规模下,对每个行进行加密评分过慢,因此我们结合两个思路——降维(缩小每个向量的维度)与分层结构(路由至小型候选集而非扫描全部内容)——在单GPU上以加密方式执行。我们在三个规模差异极大的语料库上进行评估:从约1000万张人脸图像聚类得到的222049个质心的人脸语料库(512维)、DataComp-1B(1.39×10^9个512维CLIP向量)、Deep1B(10^9个96维向量)。在DataComp-1B上,针对单个标注答案,我们达到了召回率@10为0.90;若前10个结果中存在近重复图像也计入正确(该数据为网络爬取,充满重复内容),则召回率@10为0.95,在GPU上每个加密查询的耗时约为6秒;更轻量的配置则达到0.78/0.83,耗时约1.8秒。这些是可部署的服务器端延迟,不包含客户端解密与网络传输。在Deep1B上,我们在全级别FHE下达到召回率@10为0.90(2000次FHE查询测得为0.9045,与明文路由的0.906匹配;96→128的零填充是精确的,相关性为1.0),每个查询的暖延迟为2.3秒。我们详细描述了完整的客户端-服务器协议,足以复现,并报告了所有配置的准确率与延迟。我们还测量了该速度的代价:分层结构的访问模式会泄露数据库几何信息(观察者仅从访问模式就能恢复72%的粗单元邻接图),我们证明带种子(固定分组)的填充可将该泄露降低约35倍(至约2%),而朴素填充会被重复查询攻击攻破。
英文摘要
We build a system that answers "which database vectors are most similar to my query?" without the server ever seeing the query. The query is encrypted with fully homomorphic en- cryption (FHE); the server does all its scoring on ciphertexts and returns encrypted results that only the client can read. The challenge is speed: at a billion vectors, scoring every row under encryption is far too slow, so we combine two ideas - rank reduction (shrink each vector's dimen- sion) and a hierarchy (route to a small candidate set instead of scanning everything) - executed under encryption on a single GPU. We evaluate on three corpora at very different scales: a face corpus of 222 049 centroids clustered from ~10 M face images (512-dim), DataComp-1B (1.39 x 10^9 vectors, 512-dim CLIP), and Deep1B (10^9 vectors, 96-dim). On DataComp-1B we reach a recall@10 of 0.90 against the single labeled answer, or 0.95 when a near-duplicate im- age in the top-10 also counts as correct (the data is web-scraped and full of duplicates), at ~6 s per encrypted query on a GPU; a lighter configuration reaches 0.78/0.83 at ~1.8 s. These are warm (deployable) server-side latencies - client decryption and network transfer are excluded. On Deep1B we reach recall@10 0.90 under all-levels FHE (0.9045 measured over 2000 FHE queries, matching the 0.906 plaintext routing - the 96 -> 128 zero-pad is exact, correlation 1.0) at 2.3 s warm per query. We describe the full client-server protocol in enough detail to repro- duce it, and report accuracy and latency for every configuration. We also measure what this speed costs: the hierarchy's access pattern leaks the database geometry (an observer recovers 72% of the coarse-cell neighbor graph from access patterns alone), and we show that seeded (fixed-group) padding cuts this leak by ~35x (to ~2%), where naive padding is defeated by a repeated-query attack.