arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Semord:面向分布式向量搜索的语义保持放置与低扇出路由学习

Semord: Learned Semantic-Preserving Placement and Low-Fanout Routing for Distributed Vector Search

Shengze Wang, Yi Liu, Yifan Hua, Xiaoxue Zhang, Chen Qian

arXiv 2609.25514首次发表:更新:

发表机构

University of California, Santa Cruz; University of Nevada, Reno(加州大学圣克鲁兹分校; 内华达大学里诺分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Semord提出去中心化向量搜索覆盖系统,通过VHash语义放置和VecDHT协议实现高召回低扇出路由,无需集中协调器,提升召回率15%以上并减少60%联系节点。

AI 中文摘要

向量数据库日益部署在分布式环境中,其中不同的用户、站点或域维护向量数据。现有的向量数据库依赖协调器来记录哪些分片存储向量空间的哪些部分,并将每个查询路由到这些分片。在去中心化环境中,对等节点可能加入、离开或移动数据,而没有可信节点跟踪每次更改,因此过时的路由信息可能导致查询发送到错误的节点,或需要联系许多节点,从而降低向量检索召回率并增加网络延迟。我们提出Semord,一种去中心化向量搜索覆盖系统,通过将每个ANN查询路由到一小部分相关节点来实现高召回率,而无需依赖集中式协调器。Semord通过使语义局部性可路由来解决此问题:1)我们提出VHash,将语义相关的向量放置在覆盖键空间中彼此靠近的位置,同时避免负载不均衡,以便每个查询只需联系一小部分邻居节点进行分布式局部ANN排序。2)我们设计VecDHT,一种通信协议,在成员资格和工作负载变化下维护去中心化路由、区域元数据、抗扰动性以及VHash更新。我们在真实测试平台上的大量实验表明,与去中心化基线相比,Semord将召回率提高了超过15%,并将联系的节点数减少了超过60%。Semord还接近集中式Oracle基线的召回率和延迟,同时将节点本地ANN索引峰值内存减少超过2倍。受控的大规模模拟进一步表明,Semord在真实世界嵌入工作负载上可扩展,并在扰动下作为去中心化覆盖进行范围向量检索时保持稳健。

英文摘要

Vector databases are increasingly deployed in distributed settings where different users, sites, or domains maintain vector data. Existing vector databases rely on a coordinator to record which shards store which parts of the vector space and to route each query to those shards. In a decentralized setting, peers may join, leave, or move data without a trusted node tracking every change, and outdated routing information can therefore send queries to the wrong peers or require contacting many peers, reducing vector retrieval recall and increasing network latency. We present Semord, a decentralized vector search overlay system that achieves high recall by routing each ANN query to a small set of relevant peers, without relying on a centralized coordinator. Semord addresses this problem by making semantic locality routable: 1) We propose VHash to place semantically related vectors near each other in the overlay key space while avoiding load imbalance, so that each query only needs to contact a small neighborhood of peers for distributed local ANN ranking. 2) We design VecDHT, a communication protocol that maintains decentralized routing, region metadata, churn resilience, and VHash updates under membership and workload changes. Our extensive experiments on a real testbed show that Semord improves recall by more than 15% and reduces contacted peers by over 60% compared with decentralized baselines. Semord also approaches the recall and latency of a centralized oracle baseline while reducing peak peer-local ANN index memory by more than 2X. Controlled large-scale simulations further show that Semord scales across real-world embedding workloads and remains robust under churn for scoped vector retrieval as a decentralized overlay.

Comments19 pages, 14 figures, 4 tables. Includes appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑