arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

批处理反馈与基于搜索的图构建中的随机访问瓶颈

Batched Feedback and the Random-Access Wall in Search-Based Graph Construction

Édgar Chávez

arXiv 2609.30493首次发表:更新:

发表机构

CICESE(墨西哥西南科学研究与高等教育中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究测量了基于搜索的图构建中批处理与增量构建的性能差异,发现自适应距离评估导致的随机访问瓶颈是主要限制,并提出双上限屋顶线模型来界定其影响。

AI 中文摘要

用于最近邻搜索的可导航图要么增量构建,即每个插入点在一个随构建过程而变化的图中进行搜索,要么在固定基底上批量构建,后者以构建时间为代价换来了确定性和并行性。我们测量了这一代价并找出其来源。在一台64线程机器上,我们对调优后的Vamana和PiPNN进行性能剖析,并以配对设计重复构建每个系统,将构建时间分解为工作量(距离评估次数)和每次评估的成本。让批量构建器的基底在$B$个同步块中变化,可以恢复增量构建的反馈循环,同时图仍然是(数据、种子、参数、$B$)的函数:八棵具有批处理反馈的树匹配三十二棵冻结树,并且构建完成了Vamana距离工作量的$0.88\times$。然而,它花费了Vamana墙钟时间的$1.55\times$,因为变化的基底每次评估成本更高。进一步推进,我们发现了一个任何基于搜索的构建器都无法跨越的瓶颈。它们中的每一个,无论是增量还是批处理,都自适应地评估距离,一次一个依赖的随机访问,在同一台机器上,在$d \thickapprox 100$时,其运行速度比密集内核低$20$-$27\times$,当数据驻留在缓存中时,低五到六倍。这一差距并非低维伪影:自适应评估成本为$d^{1.04}$,而分块评估成本为$d^{0.63}$,因此瓶颈随$d^{0.4}$增长,在$d = 960$时比在$d = 128$时高一倍。PiPNN的数量级构建优势在于该内核:它评估每个点的距离数量相同,但以预先固定的密集块进行评估。我们表明,波束不能事后批处理(同步块的有用密度为2-3%),将瓶颈表述为双上限的屋顶线,并界定其范围:每当距离是黑盒或评估顺序依赖于数据时,该瓶颈就会生效。

英文摘要

Navigable graphs for nearest-neighbor search are built either incrementally, each inserted point searching a graph that mutates as construction proceeds, or in batch over a fixed substrate, which buys determinism and parallelism at a price in build time. We measure that price and find where it comes from. Instrumenting a tuned Vamana and PiPNN and building every system repeatedly in a paired design on one 64-thread machine, we separate build time into work (distance evaluations) and cost per evaluation. Letting the batch builder's substrate mutate in $B$ synchronous blocks recovers the feedback loop of incremental construction while the graph stays a function of (data, seed, parameters, $B$): eight trees with batched feedback match thirty-two frozen trees, and the build does $0.88\times$ Vamana's distance work. Yet it takes $1.55\times$ Vamana's wall-clock, because the mutating substrate costs more per evaluation. Pushing further, we find a wall that no search-based builder crosses. Every one of them, incremental or batched, evaluates distances adaptively, one dependent random access at a time, and runs $20$-$27\times$ below a dense kernel on the same machine at $d \approx 100$, five to six times of it with the data resident in cache. The gap is not a low-dimension artifact: an adaptive evaluation costs $d^{1.04}$ and a blocked one $d^{0.63}$, so the wall grows as $d^{0.4}$ and is twice as high at $d = 960$ as at $d = 128$. PiPNN's order-of-magnitude build advantage is that kernel: it evaluates as many distances per point, but as fixed-in-advance dense blocks. We show the beam cannot be batched after the fact (the useful density of a lockstep block is 2-3%), state the wall as a two-ceiling roofline, and delimit it: it binds whenever the distance is a black box or the evaluation order is data-dependent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑