AI 中文总结
提出misi度量倒排样本索引,将NAPP索引推广到线性词汇表,具并行构造、低内存特性,适用于需权衡构造成本等的场景,0.99召回率验证预算随n^0.30增长。
AI 中文摘要
我们提出了misi,这是一种用于一般度量空间上近似最近邻搜索的倒排索引,其词汇表是数据库的随机样本,大小与n成比例。每个对象由其k_b个最近样本点表示,这些样本点通过样本上的可插拔内部索引找到;查询通过逆文档频率(idf)加权的共享邻居投票,随后对C个候选者进行精确验证来响应。该构造将NAPP索引从固定数量的枢轴推广到线性大小的词汇表,这使得当n增长时,倒排列表保持恒定的期望长度ρ = k_b/α,并将索引转变为一种组合:对于任何度量,αn个点上的任何高召回率索引都能生成n个点上的索引。概率模型提供了召回率保证——在重叠间隙上k_b对数级依赖于n就足够了,且验证预算由索引自身估计——以及一个匹配的极限:投票无法分辨低于1/√k_b量级的重叠差异。该设计的优势是结构性的:构造过程是n次独立搜索——极易并行化、确定性,在64核上处理10^8个向量耗时5250秒,比匹配召回率的图构建快3.7倍;它在强制3 GiB的内存限制下流式运行,且该可移植工件在强制8 GB预算下可从NVMe为10^8个向量提供服务,低于SSD图基线的工作下限。其代价是查询时的工作量:饱和图基线在RAM中回答速度快6-16倍,且0.99召回率的验证预算随n^0.30增长。所有结果都包含种子、饱和扫描和完整配置,由运行清单生成,包括测得的负面结果。预期应用场景是那些更看重构造成本、确定性、内存占用或黑盒度量而非峰值吞吐量的领域:频繁重建的语料库、批量相似性工作负载、受内存限制的服务。
英文摘要
We present misi, an inverted index for approximate nearest-neighbor search over general metric spaces whose vocabulary is a random sample of the database, of size proportional to $n$. Each object is represented by its $k_b$ nearest sample points, found by a pluggable inner index over the sample; queries are answered by an idf-weighted shared-neighbor vote followed by exact verification of $C$ candidates. The construction generalizes the NAPP index from a constant number of pivots to a linear-size vocabulary, which keeps posting lists at constant expected length $ρ= k_b/α$ as $n$ grows and turns the index into a combinator: any high-recall index on $αn$ points yields an index on $n$ points, for any metric. A probabilistic model gives a recall guarantee -- $k_b$ logarithmic in $n$ over the overlap gap suffices, with a verification budget the index itself estimates -- and a matching limit: the vote cannot resolve overlap differences below order $1/\sqrt{k_b}$. The design's strengths are structural: construction is $n$ independent searches -- embarrassingly parallel, deterministic, $5{,}250$ s for $10^8$ vectors on 64 cores, $3.7\times$ faster than a matched-recall graph build -- it streams under an enforced 3 GiB cap, and the portable artifact serves $10^8$ vectors from NVMe within an enforced 8 GB budget, below the working floor of the SSD-graph baseline. Its cost is query-time work: saturated graph baselines answer $6$-$16\times$ faster in RAM, and the verification budget for 0.99 recall grows as $n^{0.30}$. All results carry seeds, saturation sweeps and full configurations, are generated from run manifests, and include measured negative results. The intended applications weight construction cost, determinism, memory footprint, or black-box metrics over peak throughput: frequently rebuilt corpora, batch similarity workloads, constrained-memory serving.
Comments14 pages. Links to code