AI 中文总结
InSituANN是一种基于IVF的ANNS引擎,将基础向量存于主机内存、就地执行精细搜索,在单个商用GPU上实现十亿级向量搜索,大幅提升吞吐量并开源。
AI 中文摘要
对十亿级向量数据集的近似最近邻搜索(ANNS)已成为现代检索系统的基础操作,支撑着大规模推荐、语义搜索以及大语言模型/检索增强生成(LLM/RAG)工作负载。尽管GPU提供了大规模并行性和高带宽内存以用于批量向量搜索,但其有限的显存(VRAM)容量使得完全驻留于GPU的十亿级索引难以部署。在CPU-GPU异构设计中,将基础向量保留在主机内存可规避这一容量限制,但将精细搜索直接卸载到GPU会引入新的瓶颈:大量基础向量数据必须通过PCIe进行传输。我们提出InSituANN,这是一种基于IVF的ANNS引擎,可在单个商用GPU上实现十亿级向量搜索。InSituANN将原始基础向量保留在主机内存中,就地执行精细搜索,并利用GPU进行紧凑路由和可选剪枝。因此,查询处理避免了高维基础向量的PCIe传输,同时保留了IVF的简洁性。除查询性能外,我们还为InSituANN设计了超快速的IVF构建路径。在SIFT-1B数据集上,InSituANN构建IVF索引耗时5.2分钟,比实测耗时30.4小时的HNSW构建速度快约350倍。在十亿级数据集上达到匹配的召回率时,InSituANN的端到端吞吐量相比受PCIe限制的Rummy基准提升了104.9倍至4298.2倍,在SIFT-1B和DEEP-1B数据集上相比DiskANN提升了2.4倍至4.6倍。结合强大的召回率-吞吐量权衡以及比基于图的替代方案更低的索引空间,这些成果使得在高性价比硬件上实现十亿级检索成为可能。我们在该httpsURL开源了InSituANN。
英文摘要
Approximate nearest neighbor search (ANNS) over billion-scale vector datasets has become a foundational operator for modern retrieval systems, powering large-scale recommendation, semantic search, and LLM/RAG workloads. Although GPUs offer massive parallelism and high-bandwidth memory for batched vector search, their limited VRAM capacity makes fully GPU-resident billion-scale indexes difficult to deploy. In CPU-GPU heterogeneous designs, keeping the base vectors in host memory avoids this capacity limit, but naively offloading fine search to the GPU introduces a new bottleneck: large volumes of base-vector data must be streamed over PCIe. We present InSituANN, an IVF-based ANNS engine that enables billion-scale vector search on a single commodity GPU. InSituANN keeps original base vectors in host memory, performs fine search in situ, and uses the GPU for compact routing and optional pruning. As a result, query processing avoids PCIe transfers of high-dimensional base vectors while retaining the simplicity of IVF. Beyond query performance, we further design an ultra-fast IVF construction path for InSituANN. On SIFT-1B, InSituANN builds the IVF index in 5.2 minutes, about 350x faster than the measured 30.4-hour HNSW build. At matched recall on billion-scale datasets, InSituANN improves end-to-end throughput by 104.9x-4298.2x over the PCIe-bound Rummy baseline and by 2.4x-4.6x over DiskANN on SIFT-1B and DEEP-1B. Together with strong recall-throughput trade-offs and lower index space than graph-based alternatives, these gains make billion-scale retrieval practical on cost-efficient hardware. We open-source InSituANN at https://github.com/mindtravel/InSituANN-OpenSource.
Comments16 pages, 19 figures