arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于图的过滤近似最近邻搜索的简单快速算法(完整版)

Simple and Fast Algorithm for Graph-based Filtered Approximate Nearest Neighbor Search (Full Version)

Reon Uemura, Keito Kido, Daichi Amagata

arXiv 2607.24173首次发表:更新:

AI 中文总结

针对机器学习中高维向量的过滤近似最近邻搜索问题,现有技术存在性能慢、处理属性组合难的问题。本文提出新算法,经广泛实验验证了该算法在解决此问题上的高效性。

AI 中文摘要

由于基于机器学习的嵌入技术的普及,将许多对象表示为高维向量已很常见。分析高维向量的最重要功能之一是近似最近邻搜索,即给定一个查询向量,检索与查询向量近似最相似的向量。在许多实际应用中,如电子商务,对象不仅有向量,还有属性,如类别、颜色和品牌,这就需要一种用户可以指定查询向量和感兴趣的每个属性值的场景。这个问题称为过滤近似最近邻搜索,它从具有指定属性值的对象集中检索近似最近邻。有效解决这个问题具有挑战性,因为它必须接受任意查询向量和属性值,而这些事先是未知的。现有技术存在搜索性能慢和处理属性任意组合困难的问题。这项工作克服了这些挑战,提出了一种针对此问题的新算法。我们进行了广泛的实验,结果证明了我们算法的效率。

英文摘要

It has been common to represent many objects as high-dimensional vectors due to the proliferation of machine learning-based embedding techniques. One of the most important functions for analyzing high-dimensional vectors is approximate nearest neighbor search, which, given a query vector, retrieves the vector that is approximately the most similar to the query vector. In many real-world applications, such as e-commerce, objects have not only vectors but also attributes, e.g., category, color, and brand, and they require a scenario where users can specify a query vector and a value for each attribute of interest. This problem, called filtered approximate nearest neighbor search, retrieves approximate nearest neighbors from a set of objects that have the specified attribute values. Efficiently solving this problem is challenging because it has to accept arbitrary query vectors and attribute values, which are not known in advance. Existing techniques suffer from slow search performance and difficulty in dealing with arbitrary combinations of attributes. This work overcomes these challenges and proposes a new algorithm for this problem. We conduct extensive experiments, and the results demonstrate the efficiency of our algorithm.

CommentsA shorter version has been accepted to SISAP2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑