发表机构
University of North Carolina at Chapel Hill; Georgia Institute of Technology; University of Twente(北卡罗来纳大学教堂山分校; 佐治亚理工学院; 特温特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对点过程输入的分层导航,确定了保证贪婪搜索精度的几何覆盖条件,推导了固定维度下贪婪跳数的对数级期望,为分层近似最近邻搜索提供了理论依据。
AI 中文摘要
大规模信息检索系统,包括检索增强生成(RAG)和推荐引擎,广泛采用多层分层数据结构,用于高维向量空间中的超快速近似最近邻搜索。然而,确保贪婪导航准确高效的几何条件仍鲜为人知。本研究针对由d维环面𝕋ᵈ上的n个数据点构建的邻近图分层结构,研究贪婪导航的效率。我们确定了一种确定性覆盖条件,在该条件下,对于任意查询q∈𝕋ᵈ,贪婪搜索返回的点与最近点的距离在(1+ε)倍以内。当数据以齐次泊松过程、厄米特行列式过程或有界密度考克斯过程分布时,只要满足d = o(log n / log log n),该覆盖性质以高概率成立。在相同假设下,固定查询的贪婪跳数期望为O(exp(½ d log d + O(d)) log n),在固定维度下,跳数期望呈对数级增长。
英文摘要
Large-scale information retrieval systems, including retrieval-augmented generation (RAG) and recommendation engines, widely use multi-layered hierarchical data structures for ultra-fast approximate nearest-neighbor search in high-dimensional vector spaces. However, the geometric conditions that ensure accurate and efficient greedy navigation remain poorly understood. In this work, we study the efficiency of greedy navigation on a hierarchy of proximity graphs constructed from \(n\) data points on the \(d\)-dimensional torus~$\mathbb{T}^d$. We identify a deterministic coverage condition under which, given any query $q\in \mathbb{T}^d$, greedy search returns a point within $(1+\varepsilon)$-factor of the distance to the closest point. This coverage property holds with high probability when the data is distributed as a homogeneous Poisson process, a Hermitian determinantal process, or a bounded-density Cox process, as long as $d = o(\log n/\log \log n)$. Under the same assumptions, the expected number of greedy hops for a fixed query is \(O\!\left(\exp\!\left(\tfrac12 d\log d+O(d)\right)\log n\right)\), yielding logarithmic expected hop count in fixed dimension.
Comments22 pages, 7 figures