arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超高维 $\ell_\infty$ 空间中的近似最近邻

Approximate Nearest Neighbor in Ultra-High Dimensional $\ell_\infty$

Nathan White, Tian Zhang

arXiv 2609.09427首次发表:更新:

发表机构

University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对超高维 $\ell_\infty$ 空间中的近似最近邻问题,提出子集嵌入方法,实现查询时间不依赖维度,并给出匹配下界,达到 $O(1)$-近似。

AI 中文摘要

我们研究了在超高维设置下 $\ell_\infty$ 空间中的近似最近邻问题,其中维度 $d$ 显著大于点的数量 $n$。因此,我们希望数据结构在查询时间上不依赖于 $d$。[Herold-Nanongkai-Spoerhase-Varma-Wu, SoCG 2025] 引入了这个问题,并给出了 $\ell_p$ 空间中的数据结构:对于 $p=1,2$,他们给出了 $(1+\varepsilon)$-近似数据结构,空间复杂度为 $\tilde{O}(n\log d/\text{poly}(\varepsilon))$,查询时间为 $\tilde{O}(n/\text{poly}(\varepsilon))$。由于任何数据结构都必须具有 $\Omega(\min \{n,d\})$ 的查询时间,这个查询时间几乎是紧的。然而,他们的结果对于 $\ell_\infty$ 是低效的,查询时间为 $\Omega(nd)$。为了处理 $\ell_\infty$ 的挑战,我们引入了子集嵌入的概念,它通过简单地选择维度子集来嵌入点。特别地,我们表明可以通过仅计算 $n^{1+1/c}$ 个坐标上的距离来保留 $n$ 个点数据集的所有成对距离,误差因子为 $O(c)$。我们还展示了一个匹配的下界:对于任何 $c > 1$,存在 $\mathbb{R}^{d}$ 中的 $n$ 个点的集合,使得该集合的任何具有近似因子 $c$ 的子集嵌入必须至少具有 $n^{1+\Omega(1/c)}$ 个坐标。利用我们的子集嵌入,我们给出了 $\ell_\infty$ 空间中近似最近邻的数据结构,空间复杂度为 $O(n^2\log d)$,查询时间为 $\tilde{O}(n^{1+1/c})$,对于任何 $c \geq 1$,近似因子为 $O(c\log\log n)$。最后,我们给出了另一个 $\ell_\infty$ 下近似最近邻的数据结构,其空间和查询时间与我们的子集嵌入方法相同,但近似因子为 $O(c^{\log_2 3}) \approx O(c^{1.58})$。这使我们能够实现 $O(1)$-近似,例如查询时间为 $n^{1.01}$。

英文摘要

We study the approximate nearest neighbor problem under $\ell_\infty$ in the ultra-high dimensional setting where the dimension $d$ is significantly larger than the number of points $n$. Thus, we desire data structures with no dependence on $d$ in the query time. [Herold-Nanongkai-Spoerhase-Varma-Wu, SoCG 2025] introduce this problem and give data structures in $\ell_p$: for $p=1,2$, they give $(1+\varepsilon)$-approximation data structures with space $\tilde{O}(n\log d/\text{poly}(\varepsilon))$ and query time $\tilde{O}(n/\text{poly}(\varepsilon))$. Since any data structure must have query time $Ω(\min \{n,d\})$, this query time is nearly tight. However, their results are inefficient for $\ell_\infty$, with query time $Ω(nd)$. In order to handle the challenges of $\ell_\infty$, we introduce a notion of subset embeddings, which embed points by simply selecting a subset of dimensions. In particular, we show one may preserve all pairwise distances of an $n$ point dataset up to a factor of $O(c)$ by computing distances on only $n^{1+1/c}$ coordinates. We also show a matching lower bound: for any $c > 1$, there exists a set of $n$ points in $\mathbb{R}^{d}$ such that any subset embedding for the set with approximation $c$ must have at least $n^{1+Ω(1/c)}$ coordinates. Using our subset embeddings, we give data structures for approximate nearest neighbor in $\ell_\infty$ with space $O(n^2\log d)$, query time $\tilde{O}(n^{1+1/c})$, and approximation $O(c\log\log n)$ for any $c \geq 1$. Finally, we give another data structure for the approximate nearest neighbor under $\ell_\infty$ with the same space and query time as our subset embedding approach, but with approximation $O(c^{\log_2 3}) \approx O(c^{1.58})$. This allows us to achieve $O(1)$-approximation with query time e.g.~$n^{1.01}$

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑