BOA:GPU上过滤式ANNS的波束宽度在线自适应
BOA: Beamwidth Online Adaptation for Filtered-ANNS on a GPU
查看机构详情
- University of California Riverside(加州大学河滨分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出BOA,一种在GPU上通过在线波束宽度自适应,按查询定制搜索力度,在保持高召回率的同时将吞吐量提升7至12.5倍的过滤式ANNS引擎。
中文摘要 AI 辅助
过滤式近似最近邻搜索,即在满足一个或多个属性谓词的向量中返回与查询向量最近的top-k向量,已成为现代向量搜索系统中的基本操作。基于图的方法采用波束搜索来并行处理一批查询以实现高吞吐量,并使用100或更大的高固定波束宽度来确保高召回率。我们观察到,对于一批查询,在多个数据集上,超过一半的查询可以用仅50或更小的波束宽度精确求解。因此,基于固定高波束宽度的现有系统通过强制每个查询像批中最难的查询一样彻底搜索来牺牲吞吐量以实现高召回率,尽管大多数查询可以通过浅层搜索解决。在本文中,我们提出了一个名为BOA的用于单个GPU的过滤式ANNS引擎,该引擎使用在线波束宽度自适应来在多属性范围过滤器下定制批内查询的搜索力度。我们通过多阶段搜索来解决召回率与吞吐量的权衡:所有查询首先在窄波束下评估,只有结果不确定的查询才逐步使用更宽的波束宽度进行细化。这使得召回率在很大程度上对起始波束宽度不敏感,而先前的方法必须使用固定的高波束宽度才能获得高召回率。BOA+重叠各阶段的执行以进一步提高吞吐量。我们的实验表明,对于10,000个查询,在线自适应在平均波束宽度从22到77的范围内实现了94.05%至99.96%的召回率,而非自适应方法需要固定波束宽度500才能达到相似或更低的召回率。因此,自适应性将吞吐量提高了7倍至12.5倍。
英文摘要
Filtered approximate nearest neighbor search, i.e. returning the top-k vectors nearest to a query vector among those satisfying one or more attribute predicates, has become a fundamental operation in modern vector search systems. Graph-based solutions employ beam search to solve a batch of queries in parallel for high throughput and employ high fixed beamwidth of 100 or greater for ensuring high recall. We observe that, given a batch of queries, more than half of the queries across multiple data sets can be solved precisely with a beamwidth of just 50 or less. Therefore, existing systems based on fixed high beamwidth sacrifice throughput to achieve high recall by forcing every query to search as thoroughly as the hardest query in the batch even though majority of queries can be resolved by a shallow search. In this paper we present a filtered ANNS engine for a single GPU named BOA that uses online beamwidth adaptation to customize the search effort across queries within a batch under multi-attribute range filters. We address the recall throughput tradeoff with a multi-phase search: all queries are first evaluated under a narrow beam, and only those with uncertain results are progressively refined with wider beamwidths. This renders recall largely insensitive to the starting beamwidth, whereas prior methods must use a fixed high beamwidth for high recall. BOA+ overlaps execution of phases to further enhance throughput. Our experiments show that, for 10,000 queries, online adaptation achieves 94.05% to 99.96% recall with average beamwidth ranging from 22 to 77, while a non-adaptive approach requires a fixed beamwidth of 500 to achieve similar or lower recall. Consequently, adaptivity increases throughput by 7x to 12.5x