arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21494cs.IRcs.DBcs.LG

检索需要多向量:指数级分离

Retrieval Needs Multivectors: An Exponential Separation

  • Microsoft Research India(微软研究院印度分部)

机构由 AI 辅助整理,请以论文原文为准。

Mihir Agarwal, Viraj Agrawal, Sabyasachi Basu, Ankit Garg, Kirankumar Shiragur

AI总结:

本研究通过理论构造确立了文档排序任务中单、多向量嵌入表达能力的指数级分离,引入检索基准ANDOR并验证多向量模型在该基准上的性能优势。

AI中文摘要:

近期研究通过理论分析和LIMIT等具有挑战性的基准,凸显了基于嵌入的检索模型的表达能力局限。虽然多向量嵌入始终优于单向量嵌入,但两者间的确切表达差距仍未被充分理解。本研究遵循Jayaram的工作,首次提供了明确的查询与文档集合及其相关性矩阵,对于该集合,能将所有相关文档排在不相关文档之上的单向量嵌入需要指数级大小,而多项式大小的多向量嵌入即可满足要求。我们的结果确立了文档排序任务中,单向量与多向量嵌入表达能力的指数级分离,这与Jayaram工作中近似数值分数的情况形成对比。受理论构造启发,我们引入了新的检索基准ANDOR,其自然实例化了这些困难示例。我们发现,在零样本设置下,最先进的单向量嵌入模型在ANDOR上表现不佳,微调后仅显示出微小提升,凸显了该基准相较于以往工作的固有难度。相比之下,多向量模型始终优于其单向量对应模型,且微调后大幅提升,与我们的理论预测高度一致。

英文摘要:

Recent works have highlighted the expressive limitations of embedding based retrieval models through both theoretical analyses and challenging benchmarks such as LIMIT. While multi-vector embeddings consistently outperform single-vector embeddings, the precise representational gap between them remains poorly understood. In this work, following Jayaram's work, we provide the first explicit family of query and document sets, together with their relevance matrices, for which single-vector embeddings that rank all relevant documents above irrelevant ones require exponential size, whereas polynomial-size multi-vector embeddings suffice. Our result establishes an exponential separation between the expressive power of single-vector and multi-vector embeddings for the task of ranking of documents as opposed to approximating numerical scores as in the work of Jayaram. Motivated by our theoretical construction, we introduce ANDOR, a new retrieval benchmark that naturally instantiates these hard examples. We show that state-of-the-art single-vector embedding models perform poorly on ANDOR in the zero-shot setting and exhibit only marginal improvements after fine-tuning, highlighting the inherent difficulty of the benchmark compared to prior work. In contrast, multi-vector models consistently outperform their single-vector counterparts and improve substantially with fine-tuning, closely aligning with our theoretical predictions.

↑