AI 中文总结
本研究针对外包加密向量数据库的隐私保护需求,提出了分离范围定位与加密向量搜索的过滤-精调RFANNS方案,实验验证其在QPS-召回率权衡上优于现有安全适配方案。
AI 中文摘要
范围过滤近似最近邻搜索(RFANNS)是向量数据库的重要基础操作,用于检索与查询向量相似且满足数值范围谓词的向量,但现有RFANNS索引会以明文形式暴露向量、属性和查询。该假设不适用于外包向量数据库,此类场景中敏感数据和查询必须免受诚实但好奇的云服务器的侵害。据我们所知,本研究首次系统地提出并评估了外包加密向量数据库上的隐私保护RFANNS方案。我们的方法将范围定位与加密向量搜索分离:授权用户将查询范围映射到本地N叉属性树的紧凑节点集合,服务器仅在加密向量上搜索对应的邻近图子索引。为减少开销较大的加密比较,我们采用过滤-精调流水线:首先通过近似距离比较保持加密检索粗候选,再通过精确距离比较加密对小候选集重排序。随后我们分析了该协议的计算、存储、通信及泄漏情况。在四个广泛使用的向量数据集上的实验表明,与现有RFANNS方法的代表性安全适配方案相比,我们的方法优化了QPS-召回率的权衡,可有效扩展至大型数据集。
英文摘要
Range-filtered approximate nearest neighbor search (RFANNS) is an important primitive for vector databases; it retrieves vectors that are similar to a query and satisfy a numerical range predicate, but existing RFANNS indexes expose vectors, attributes, and queries in plaintext. This assumption is unsuitable for outsourced vector databases, where sensitive data and queries must be protected from an honest-but-curious cloud server. To the best of our knowledge, this is the first study that systematically formulates and evaluates privacy-preserving RFANNS over outsourced encrypted vector databases. Our approach separates range localization from encrypted vector search: an authorized user maps the query range to a compact set of nodes in a local N-ary attribute tree, and the server searches only the corresponding proximity graph sub-indices over encrypted vectors. To reduce expensive encrypted comparisons, we use a filter-and-refine pipeline that first retrieves coarse candidates with approximate distance-comparison-preserving encryption and then reranks a small candidate set with exact distance-comparison encryption. We then analyze the computation, storage, communication, and leakage of the protocol. Experiments on four widely used vector datasets show that our method improves the QPS-Recall trade-off over representative secure adaptations of existing RFANNS approaches, scaling effectively to large datasets.
CommentsAccording to the best of our knowledge, this work is the first attempt to study privacy-preserving range-filterd ANN search problem. This is the early version of the work that is still in progress