隐藏在对数外衣下的幂律:基于图的向量搜索的可扩展性研究
A Power Law in Logarithm's Clothing: On the Scalability of Graph-Based Vector Search
浏览论文内容
中文总结 AI 辅助
该研究针对基于图的向量搜索的可扩展性,发现小数据集规模下搜索成本呈次线性幂律增长、大数据集下转为亚多项式增长,提出统一理论与权衡模型,为向量数据库优化提供依据。
中文摘要 AI 辅助
大多数向量数据库依赖基于图的索引(尤其是HNSW和Vamana)进行近似最近邻搜索。随着嵌入模型被广泛采用,这些数据库存储的数据集规模迅速增长。在固定准确率下,搜索成本如何随数据集规模变化?主流答案是对数多项式增长。然而,这一结论仅在特定条件下被证明,对于实际使用的索引则未经证明,且基本未经过测试:标准基准仅在单一数据集规模下测量成本,而非跨多个规模。我们对该结论进行了验证,发现答案取决于规模本身:当数据集规模N相对于数据的内在维度较小时,搜索成本按N^c增长,其中常数0<c<1,我们将这种缩放称为次线性幂律;一旦N足够大,增长会放缓至亚多项式,与对数多项式的结论一致。在我们测试的每个数据集、几乎所有数据集规模、每个召回率目标、查询难度级别和索引配置下,次线性幂律均会出现;而当两个数据集的规模相对于其内在维度足够大时,会出现向亚多项式增长的转变。两种行为背后的机制是:数据集的内在维度随其规模增长,直到数据解析出其潜在分布;更高的内在维度会将更多向量打包到搜索必须检查的查询邻域中。我们提出了一种统一的束搜索成本理论来解释我们的观察结果:对于精确和有界度的构造,我们证明了次线性幂律及其向对数多项式缩放的最终转变,并推导了该转变发生的规模;我们还开发了可预测任意召回率目标和索引配置的幂律指数的模型,这些模型为数据增长时在搜索成本、插入成本和召回率之间进行权衡提供了原则性方法。
英文摘要
Most vector databases rely on graph-based indexes, notably HNSW and Vamana, for approximate nearest neighbor search. With embedding models widely adopted, the datasets these databases store grow rapidly. At a fixed accuracy, how does search cost scale with dataset size? The prevailing answer is poly-logarithmic growth. Yet the claim is proven only under special conditions and asserted without proof for the indexes used in practice. It is also largely untested: standard benchmarks measure cost at one dataset size, not across sizes. We put the claim to the test. The answer depends on the scale itself. While the dataset size $N$ is small relative to the data's intrinsic dimensionality, search cost grows as $N^c$ for a constant $0<c<1$. We call this scaling the Sublinear Power Law. Once $N$ is large enough, growth slows to subpolynomial, consistent with the poly-logarithmic claim. The Sublinear Power Law appears on every dataset, mostly up to its full size, at every recall target, query hardness level, and index configuration we test. The transition to subpolynomial growth appears on the two datasets that grow large enough relative to their intrinsic dimensionality. One mechanism underlies both behaviors: a dataset's intrinsic dimensionality grows with its size until the data resolves its underlying distribution. Higher intrinsic dimensionality packs more vectors into the query neighborhood the search must examine. We present a unifying theory of beam-search cost that explains our observations. For exact and bounded-degree constructions, we prove the Sublinear Power Law and the eventual transition to poly-logarithmic scaling, and derive the scale at which it occurs. We also develop models that predict the power-law exponents for any recall target and index configuration. These models give a principled way to navigate trade-offs among search cost, insertion cost, and recall as data grows.
发表机构
- University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。