发表机构
Temple; Microsoft Research India(天普大学; 微软印度研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EigenLI利用晚期交互表示的低秩结构,通过文档特定子空间的谱近似压缩表示,并衍生出单向量表示EigenLI-SV,在多个模型和数据集上优于聚类方法和MUVERA。
AI 中文摘要
晚期交互模型(如ColBERT)通过用多个词级向量表示每个文档来实现强大的效果,但这种表达性导致了高昂的索引成本、存储占用和昂贵的MaxSim评分。我们表明,晚期交互表示具有内在的低秩结构:文档词嵌入集中在一个保留大部分检索信号的低维子空间中。利用这一观察,我们引入了EigenLI,一种谱近似框架,通过文档特定的低维子空间压缩晚期交互表示。与聚类或池化方法不同,EigenLI识别每个文档的主导特征方向,并使用它们构建缩减的交互表示。实验上,$k$-EigenLI在$k \le 32$时在ColBERTv2和AnswerAI-ColBERT-small上优于基于k-means和Ward聚类的池化方法;GTE-ModernColBERT在$k=32$时表现出不同的权衡,其中聚类方法表现更好。相同的谱构造也产生了EigenLI-SV,一种从缩减结构的二阶摘要导出的兼容ANN的单向量表示。在多个数据集和所有三种文本模型上,EigenLI-SV始终优于类似的单向量替代方法(如MUVERA)。
英文摘要
Late-interaction models such as ColBERT achieve strong effectiveness by representing each document with many token-level vectors, but this expressivity leads to large indexing cost, storage footprints and expensive MaxSim scoring. We show that late-interaction representations exhibit an intrinsic low-rank structure: document token embeddings concentrate in a low-dimensional subspace that preserves most of the retrieval signal. Leveraging this observation, we introduce EigenLI, a spectral approximation framework that compresses late-interaction representations via document-specific low-dimensional subspaces. Unlike clustering or pooling methods, EigenLI identifies the dominant eigendirections of each document and uses them to construct reduced interaction representations. Empirically, $k$-EigenLI with $k \le 32$ outperforms k-means and Ward clustering based pooling methods on ColBERTv2 and AnswerAI-ColBERT-small; GTE-ModernColBERT exhibits a different tradeoff at $k=32$, where clustering methods perform better. The same spectral construction also yields EigenLI-SV, an ANN-compatible single-vector representation derived from the second-order summary of the reduced structure. Across multiple datasets and all three text models, EigenLI-SV consistently outperforms comparable single-vector surrogates such as MUVERA.