发表机构
Stanford University; Amazon AWS; Tel Aviv University(斯坦福大学; 亚马逊云科技; 特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨向量检索中查询编码器的可学习性,发现其质量远低于文档索引的几何容量上限,并理论证明学习查询编码器在计算上可能困难,需指数级统计查询。
AI 中文摘要
高效的向量检索需要同时具备两个条件:一是语料库几何结构能够支持通过向量相似度检索到正确的文档,二是查询编码器能够将查询嵌入到其目标文档附近的嵌入空间中。近期研究从实现n个文档的所有top-k答案集所需的最小嵌入维度的角度,研究了几何容量。我们研究了几何容量的另一种概念——冻结文档索引可实现的最大召回率——并探讨学习得到的查询编码器能否达到这一上限。在多个真实世界的检索基准上,我们展示了单向量查询编码器的检索质量往往远低于文档索引所能支持的水平。受此观察启发,我们给出了理论证据表明学习查询编码器在计算上可能是困难的。特别地,我们构造了一个检索任务,该任务(1)存在一个具有完美召回率的查询编码器,该编码器可由一个小的单隐层ReLU网络表示,但(2)任何统计查询学习器(一类通过聚合统计量访问训练数据的学习器)要实现对随机基线k/n的非平凡召回优势,被证明需要指数级数量的统计查询。综合来看,我们的结果表明检索基准中存在大量未实现的几何容量,并确立了查询编码器的可学习性作为基于嵌入的检索中可能存在的障碍。
英文摘要
Efficient vector retrieval requires both a corpus geometry that supports retrieving the right documents through vector similarity, and a query encoder that can embed queries near their desired documents in the embedding space. Recent work has studied geometric capacity through the lens of the minimum embedding dimension needed to realize all top-$k$ answer sets of $n$ documents. We study a different notion of geometric capacity--the maximum recall achievable for a frozen document index--and explore whether learned query encoders can reach this ceiling. On several real-world retrieval benchmarks, we show that retrieval quality of single-vector query encoders often lies far below what the document indices can support. Motivated by this observation, we give theoretical evidence that learning query encoders can be computationally hard. In particular, we construct a retrieval task that (1) admits a query encoder with perfect recall which is representable by a small one-hidden-layer ReLU network, but (2) any statistical-query learner (a class capturing learners that access training data through aggregate statistics) provably requires exponentially many statistical queries to achieve non-trivial recall advantage over the random baseline $k/n$. Taken together, our results suggest substantial unrealized geometric capacity in retrieval benchmarks and establish query encoder learnability as a possible barrier in embedding-based retrieval.