arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03651cs.AIcs.IR

MRVQ:一个用于维度和速率弹性向量检索的驻留索引

MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

Sean Culatana, Shang-En Huang, Kang Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对稠密检索中维度与速率切换的内存开销问题,提出 MRVQ 残差量化器,通过单一驻留索引支持多种维度与速率组合,大幅降低内存占用,虽略损质量但优于多种基线。

中文摘要 AI 辅助

稠密检索服务必须在延迟、质量和内存预算变化时,在嵌入前缀维度和索引比特率之间进行切换。为每个速率单独调整量化器可获得最佳质量,但检索层随后需同时持有多个码流和量化器状态。我们引入了 Matryoshka 残差向量量化(MRVQ),一种针对冻结嵌入的后置残差量化器。其最大速率码可通过两种方式截断:丢弃残差阶段可降低速率,丢弃嵌入坐标可降低维度。因此,一个驻留产物即可服务于我们评估的所有(维度,速率)组合。在 FiQA 和 NFCorpus 上,跨四个嵌入族和 {4, 8, 16} 字节码,MRVQ 是我们评估中内存占用最低的设计。它比三个单独训练的 QINCo2 索引少使用 17.8-22.0 倍的内存,比精简的共享模型 steelman 少 1.89-2.02 倍。这种节省并非免费:每速率 QINCo2 在 FiQA 上 nDCG@10 高 0.026-0.107。但 MRVQ 在匹配码大小下优于 PQ、OPQ 和 AdANNS-OPQ。我们还评估了一种低构建成本的 PCA-标量设计,其质量可与 RaBitQ 及其扩展相媲美,同时中位数拟合速度快 420 倍。最后,我们报告两个负面结果:QINCo2 在高速率训练时崩溃,以及一个排序边界假设未达到其预定的接受标准。因此,MRVQ 是弹性检索的低内存操作点,而非普适的质量赢家。

英文摘要

Dense-retrieval services must switch among embedding-prefix dimensions and index bit rates as latency, quality, and memory budgets change. Tuning a quantizer separately for each rate gives the best quality, but the retrieval tier then holds several code streams and quantizer states at once. We introduce Matryoshka Residual Vector Quantization (MRVQ), a post-hoc residual quantizer for frozen embeddings. Its maximum-rate code can be truncated two ways: dropping residual stages lowers the rate, and dropping embedding coordinates lowers the dimension. One resident artifact therefore serves every (dimension, rate) pair we evaluate. Across FiQA and NFCorpus, four embedding families, and {4, 8, 16}-byte codes, MRVQ is the lowest-RAM design we evaluate. It uses 17.8-22.0x less memory than three separately trained QINCo2 indices, and 1.89-2.02x less than a lean shared-model steelman. The saving is not free: per-rate QINCo2 is 0.026-0.107 nDCG@10 better on FiQA. But MRVQ beats PQ, OPQ, and AdANNS-OPQ at matched code size. We also evaluate a low-build-cost PCA-scalar design that attains quality comparable to RaBitQ and its extension while fitting 420x faster at the median. Finally, we report two negative results: QINCo2 collapses when trained at high rates, and a ranking-bound hypothesis misses its pre-specified acceptance criteria. MRVQ is therefore a low-memory operating point for elastic retrieval, not a universal quality winner.

发表机构

  • Atlassian(阿特拉斯公司)
  • National Taiwan University(国立台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

↑