arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于学习型稀疏检索的紧凑高效索引

Compact and Efficient Indexes for Learned Sparse Retrieval

Franco Maria Nardini, Luca Rizzo, Cosimo Rulli, Rossano Venturini

arXiv 2610.12300首次发表:更新:

发表机构

ISTI–CNR; University of Pisa; Linkup 1(ISTI-CNR; 比萨大学; Linkup 1)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文基于SEISMIC优化倒排与正向索引,提出紧凑高效的学习型稀疏检索索引方案,在MS MARCO上的实验显示其速度与内存占用均优于现有方案。

AI 中文摘要

本文研究如何在不牺牲最先进检索数据结构效率的前提下,大幅降低学习型稀疏检索索引的内存占用。基于SEISMIC,我们重新审视其设计的两个层面:用于选择候选的倒排索引和用于评分的正向索引。对于倒排索引,我们用medoids(即被选为块代表的现有文档)替代成本高昂的每块摘要,将每块元数据从稀疏向量压缩为单个文档标识符。对于正向索引,我们对其组件和值均进行压缩:我们重新排列词汇表,使共现组件更靠近,并使用DOTPACKING8(一种SIMD友好的位打包方案,将解压与点积评估融合)对得到的Δ间隙进行编码;值则采用适配各组件分布的紧凑每组件4比特码本进行量化。我们还引入了JUMPDOT,一种专为仅含少量非零项的查询设计的分块点积核。我们的正向索引压缩独立于SEISMIC,可插入任何依赖正向索引评分的系统,我们通过将其集成到KANNOLO中对此进行了验证。在MS MARCO上使用三种最先进的学习型稀疏编码器进行的综合评估表明,我们的方案显著改善了学习型稀疏检索的速度-空间权衡:在精度相同时,我们的索引比最佳竞争对手的查询响应速度快达5.3倍,同时内存使用量减少约3倍;在内存最受限的场景中,它们仍能快达1.9倍,同时内存使用量减少达3.9倍。

英文摘要

This paper investigates how to substantially reduce the memory footprint of learned sparse retrieval indexes without sacrificing the efficiency of state-of-the-art retrieval data structures. Building on SEISMIC, we revisit both levels of its design: the inverted index used to select candidates and the forward index used to score them. For the inverted index, we replace costly per-block summaries with medoids, namely existing documents elected as block representatives, collapsing the per-block metadata from a sparse vector to a single document identifier. For the forward index, we compress both components and values. We reorder the vocabulary to place co-occurring components closer together and encode the resulting $Δ$-gaps with DOTPACKING8, a SIMD-friendly bit-packing scheme that fuses decompression with dot-product evaluation; values are quantized with compact per-component 4-bit codebooks fitted to each component's distribution. We further introduce JUMPDOT, a blocked dot-product kernel tailored for queries that contain only a few non-zero entries. Our forward-index compression is independent of SEISMIC and can be plugged into any system relying on forward-index-based scoring, as we demonstrate by integrating it into KANNOLO. A comprehensive evaluation on MS MARCO with three state-of-the-art learned sparse encoders shows that our solutions markedly improve the speed-space trade-off of learned sparse retrieval: at equal accuracy, our indexes answer queries up to 5.3x faster than the best competitor while using about 3x less memory, and in the most memory-constrained regime, they remain up to 1.9x faster while using up to 3.9x less memory.

Comments15 pages, 2 figures. Accepted at IEEE International Conference on Data Engineering 2027 (IEEE ICDE 2027)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑