发表机构
Computer Science Department, CICESE(墨西哥西南科学研究中心计算机系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SOLO提出一种无排序启发式的度量空间近似最近邻索引,通过采样倒排列表和仅扫描叶子实现可认证召回率,并在Deep-100M和Deep-1B上达到高召回与低内存,吞吐量优于HNSW。
AI 中文摘要
我们提出SOLO,一种用于一般度量空间中近似最近邻搜索的索引,其服务路径不包含任何类型的排序启发式:查询被路由到数据库随机样本的$k_s$个最近点,并且触及的发布列表中的每个对象都用真实距离进行评估。因为没有任何对象需要被排序,召回率等于一个可从存储索引计算出的覆盖概率:对查询样本进行一次真实值遍历即可同时认证所有操作点,而无需服务其中任何一个——这就是召回率证书,而对于可导航图,无论付出何种代价都不存在类似对象。整个索引是一个递归规则——对集合进行采样,将每个对象发布到其$b$个最近样本点,拆分任何超出界限的列表,始终扫描叶子节点——其操作面遵循等功定律,召回率$\approx f(b \cdot k_s)$,其水平是数据集的一个单标量特征。同样的仅扫描结构提供了一个服务下限,一旦路由器本身由同一规则索引,任何图架构都无法达到:Deep-100M在1 GB常驻内存(强制上限,每对象10.7字节)下以召回率0.9977服务,在256 MB下以0.9964服务,Deep-1B在512 MB下以召回率0.9925服务(在深度3时从96 MB),插入是一次搜索,删除是精确的。吞吐量在硬件允许的情况下具有竞争力——在双插槽32核服务器上,$10^8$规模下最高可达调优HNSW的$1.8\times$,操作点位于该图饱和点的右侧——表格在相同硬件和真实值下报告了与HNSW、DiskANN、GRAFT、NAPP、misi和SPANN分配规则的对比。
英文摘要
In every fast nearest-neighbor index, recall is measured, never predicted: each operating point is tuned by serving it against ground truth. SOLO is an index whose recall is computed from the index itself, before any query is served. SOLO is an inverted file whose vocabulary is a random sample of the database. Each object's $k_b$ nearest sample points are stored once, ranked; at serve time an object is posted under the first $b \le k_b$ of them, and a query scans the lists of its $k_s$ nearest sample points exhaustively, with the true distance. Recall depends on the product $b \cdot k_s$ (an equal-work law), so the search-side $k_s$ compensates for a small $b$ with no rebuild. There is no beam, vote or pruning bound, so a true neighbor is missed only if it shares no sample point with the query -- a membership event decided by stored integers, not by a search. Recall is therefore a count: one ground-truth pass over a sample of the operator's queries certifies every $(b, k_s)$ at once, with nothing served. The certificate matches served recall to four decimals from $10^6$ to $10^9$ objects, including 768-dimensional text embeddings under inner product with shifted queries. No graph index has an analogous object. The rest is the same rule applied recursively: a list that outgrows a bound is sampled and split like the database, and so is the vocabulary itself, which is what drives resident memory down. Deep-100M is served at recall 0.9977 from 1 GB of enforced resident memory, Deep-1B at 0.9925 from 96 MB. Inserts are one search, deletes are exact, and throughput reaches $1.8\times$ a tuned HNSW at $10^8$. Every number is reported against HNSW, DiskANN, GRAFT, NAPP, misi, SPANN, ScaNN and RaBitQ on the same hardware and ground truth.
Comments31 pages. v2: journal version. Rewritten for readability; adds a self-contained explanation of the recall certificate , the resident-memory law as its own section, ScaNN and RaBitQ arms in the capped table, and the 8-bit router served end to end at 10^9. Code and manifests: https://github.com/zevahcle/SOLO