发表机构
The University of Osaka(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有带多属性范围过滤的近似最近邻搜索(ANNS)技术的局限,提出新型框架Grant,其搜索时间与对象数量呈次线性关系,经实验验证性能优于现有技术。
AI 中文摘要
近年来,学术界和工业界开始关注带过滤条件的近似最近邻搜索(ANNS)问题。在该场景中,每个对象由高维向量和属性值组成。给定查询向量、属性值过滤条件以及k,该问题需要从满足过滤条件的对象集合中,返回与查询向量近似最近的k个向量。本文研究的是范围过滤,即用户可对每个属性指定一个范围约束。现有多数研究未考虑该场景,它们假设仅存在单个属性,或要求过滤条件需匹配相同的属性值或类别。针对这些假设的现有技术无法适用于本文场景,或效率极低。尽管预过滤、后过滤等标准解决方案可处理该问题,但效率同样低下。部分研究虽解决了与本文相同的问题,但其技术需要依赖历史查询工作负载,这极大限制了实际应用场景。为消除这些限制,本文提出了Grant,一种新型框架,可高效解决该问题,同时支持任意范围过滤和ANNS数据结构。Grant可保证搜索时间与对象数量呈次线性关系,这是现有技术无法做到的。我们开展了大量实验,结果表明Grant的性能优于现有技术。
英文摘要
Recently, academia and industry have considered the problem of approximate nearest neighbor search (ANNS) with filters. In this setting, each object consists of a high-dimensional vector and attribute values. Given a query vector, filters for attribute values, and $k$, this problem retrieves $k$ vectors approximately nearest to the query vector among a set of objects passing the filters. This paper considers range filters, i.e., users can specify a range constraint for each attribute. Most existing works do not consider this setting, and they assume (i) only a single attribute or (ii) matching filters that require the same attribute values or categories. Existing techniques for these assumptions are not available for our setting or are trivially not efficient. Although standard solutions, such as pre-filter and post-filter, can handle our problem, they are also inefficient. Some works tackle the same problem as ours, but their techniques necessitate historical query workloads, which significantly limit practical use cases. To remove these limitations, this work proposes Grant, a novel framework that solves this problem efficiently while accepting arbitrary range filters and ANNS data structures. Grant can guarantee a search time sub-linear to the number of objects, which is not held by existing techniques. We conduct extensive experiments, and their results demonstrate that Grant outperforms existing techniques.