发表机构
University of Chicago; Argonne National Laboratory; University of Wisconsin–Madison(芝加哥大学; 阿贡国家实验室; 威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
COMPASS利用知识图谱确定数据放置和查询路由,在生物医学数据上仅搜索13-18%语料库,保持召回率并提升多跳证据恢复,实现可扩展向量搜索。
AI 中文摘要
向量数据库使用哈希将数据分区到“分片”(逻辑单元)上以进行分布式执行。然而,这种放置方式破坏了语义局部性,迫使每个查询进行受限于最慢分片的散射-收集操作。向量空间聚类可以提供帮助,但科学证据通常通过事实关系相连,这些关系与嵌入距离并不对齐。我们提出了COMPASS,一个利用知识图谱(KG)确定数据放置和查询时分片选择的框架。COMPASS检测社区,拆分过大的社区,按主体实体插入嵌入,并将查询路由到一小部分分片。在四个生物医学知识图谱上,我们的方法仅搜索语料库的13-18%,同时保持广播召回率,并比基于嵌入的基线多恢复高达2.6倍的多跳证据。在15个HPC节点上,COMPASS比基于哈希的广播维持了7.9倍更高的吞吐量,且尾部延迟更低。这些结果表明,知识图谱结构为嵌入几何提供了紧凑的补充,以实现可扩展的向量搜索。
英文摘要
Vector databases use hashing to partition data across "shards," logical units for distributed execution. This placement, however, destroys semantic locality, forcing each query into scatter-gather limited by the slowest shard. Vector-space clustering can help, but scientific evidence is often connected by factual relations that do not align with embedding distance. We present COMPASS, a framework that uses a knowledge graph (KG) to determine data placement and query-time shard selection. COMPASS detects communities, splits oversized communities, inserts embeddings by subject entity, and routes queries to a small set of shards. Across four biomedical KGs, our method searches only 13-18% of the corpus while preserving broadcast recall and recovering up to 2.6x more multi-hop evidence than an embedding-based baseline. On 15 HPC nodes, COMPASS sustains 7.9x higher throughput with lower tail latency than hash-based broadcast. These results show that KG structure provides a compact complement to embedding geometry for scalable vector search.