arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

闭式解与合成孪生体:从嵌入统计量预测近似最近邻召回率

Closed Forms and Synthetic Twins: Predicting Approximate Nearest Neighbor Recall from Embedding Statistics

Shmuel Herman

arXiv 2609.00364首次发表:更新:

AI 中文总结

该研究提出从嵌入统计量预测近似索引召回率的方法,利用闭式解、合成孪生体等工具,可在索引构建前预测召回率,还能通过优化得分边际提升多索引系列的召回率,适用于动态语料库。

AI 中文摘要

嵌入模型的训练与评估均以精确检索为前提,但在生产环境中,它们需依托近似索引(如HNSW、IVF、乘积量化或晚期交互模型的固定维度编码FDE)运行,而编码器的基准测试从未涉及这些索引的行为:某现代编码器通过其原始FDE索引仅能恢复其精确前10个结果中的14%。此类失败仅在索引构建后才会显现,而标准的补救措施(如白化等适配语料库的变换)需进行拟合、存储,并随语料库变化重新拟合,还可能悄悄改写编码器的输出。本文表明,索引行为可在构建任何内容之前,通过原始嵌入的无标签统计量预测,借助与各索引系列需求匹配的一系列工具实现:(1)针对固定网格量化器(PQ、FDE)的闭式矩统计量;(2)针对分区索引的合成孪生语料库(生成的聚类统计量)上的模拟,可在该孪生语料库上构建并测试任何索引,包括生产系统;(3)针对百万文档规模图索引的尺寸外推、轻量校准孪生体。在未见过的百万文档语料库上,预测值与实测召回率的偏差在0.03以内。该几何结构可训练:针对事后变换无法改变的单一统计量——得分边际进行优化,可同时提升所有索引系列的召回率,仅产生微小的实测任务损失。最终成果包括:从一次廉价测量过程即可对新语料库和新索引进行索引选择、校正定价及生产召回率预测;无需适配语料库的变换机制,适用于持续变化的语料库;通过使编码器适配几何而非耦合到任何单一索引,推动了召回率-计算前沿的发展。

英文摘要

Embedding models are trained and evaluated as if retrieval were exact; in production they serve behind approximate indexes -- HNSW, IVF, product quantization, or the fixed-dimensional encodings (FDEs) of late-interaction models -- whose behavior the encoder's benchmarks never see: one modern encoder recovers just 14% of its exact top-10 through its raw FDE index. Such failures surface only after an index is built, and the standard patches -- corpus-fitted transforms such as whitening -- must be fitted, stored, and refit as the corpus changes, and can silently rewrite what the encoder returns. This paper shows that index behavior is predictable before anything is built, from label-free statistics of the raw embeddings, through a ladder of instruments matched to what each index family consumes: (1) closed-form moment statistics for the fixed-grid quantizers (PQ, FDE); (2) simulation on a synthetic twin corpus -- cluster statistics made generative, on which any index, composed production systems included, can be built and tested -- for partition indexes; (3) size-extrapolated, lightly calibrated twins for graph indexes at million-document scale. Predictions land within 0.03 of measured recall on an unseen million-document corpus. The same geometry is trainable: targeting the one statistic no post-hoc transform can move -- the score margin -- lifts recall for every index family at once, at a small measured task cost. The result: index choice, correction pricing, and production recall forecast from one cheap measurement pass, on new corpora and new indexes alike; serving without per-corpus transform machinery, suited to continuously changing corpora; and a recall-compute frontier pushed by adapting encoders to geometry rather than coupling them to any single index.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑