arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15058cs.DB

快速标签过滤近似最近邻搜索:通过渐进式标签集分层

Fast Label-Filtering Approximate Nearest Neighbor Search via Progressive Label Set Stratification

  • Nanjing University(南京大学)
  • State Key Laboratory for Novel Software Technology(软件新技术国家重点实验室)
  • National Institute of Healthcare Data Science(国家健康医疗大数据研究院)

机构由 AI 辅助整理,请以论文原文为准。

Ziqi Wang, Jingzhe Zhang, Shuo Shen, Wei Hu

AI总结:

本文提出标签分层相似度图(LSSG)用于标签过滤近似最近邻搜索,通过增量插入和MinHash实现高效剪枝与扩展,在相等、包含和重叠查询中显著提升速度并保持精度。

AI中文摘要:

近似最近邻搜索(ANNS)在高维空间中检索与查询向量最相似的向量。标签过滤ANNS(LFANNS)通过标签过滤器扩展了ANNS,要求基础向量的标签必须与查询标签满足某种集合关系(例如,相等、包含或重叠)。现有的LFANNS索引在不同过滤类型下性能不一致,并且在标签规模和分布变化时扩展性下降。在本文中,我们定义了标签分层相似度图(LSSG),其中边连接标签集落在分层相似度阈值内的相邻向量。为了高效实现LSSG,我们设计了一种增量插入算法,在向量和标签空间中剪除冗余边,并利用MinHash结构确保大规模标签的可扩展性。我们在显式标签模型下分析了逐步概率,并解释了为什么更严格的标签层级会减少无效的过滤内扩展。基准实验表明,LSSG在相等查询中实现了理想的最优性,在包含和重叠查询中,查询速度分别比最佳竞争索引快1.06倍至92.9倍和1.08倍至84.1倍,同时保持相同的精度,索引大小仅为0.35倍。

英文摘要:

Approximate nearest neighbor search (ANNS) retrieves the most similar vectors to a query vector in high-dimensional space. Label-filtering ANNS (LFANNS) extends ANNS with a label filter that the labels of base vectors must satisfy a set relation (e.g., equality, containment, or overlap) with the query labels. Existing LFANNS indices suffer from inconsistent performance across different filter types and degraded scalability under varying label scale and distribution. In this paper, we define label-stratified similarity graph (LSSG), where edges connect neighboring vectors whose label sets fall within stratified similarity thresholds. To implement LSSG efficiently, we design an incremental insertion algorithm to prune redundant edges in both vector and label spaces, and leverage a MinHash structure to ensure scalability for large-scale labels. We analyze stepwise probabilities under explicit label models and explain why stricter label tiers reduce ineffective in-filtering expansions. Benchmark experiments show that LSSG achieves ideal optimality for equality queries, and 1.06x-92.9x and 1.08x-84.1x faster than the best competing index for containment and overlap, respectively, in query speed with identical accuracy and 0.35x index size.

补充信息

↑