arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07276cs.IRcs.AIcs.LG

马特罗什卡哈希表示用于模型感知的紧凑语义检索

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

Peichun Hua, Yunming Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

针对RAG检索中量化码在多字节预算下性能冲突的问题,提出两阶段马特罗什卡哈希表示(MHR),分离全宽训练与前缀组织,在MS MARCO和BEIR上以32字节超越基线。

中文摘要 AI 辅助

检索增强生成(RAG)依赖于稠密检索:每个文档被存储为一个学习得到的向量,查询通过在该向量空间中找到其最近邻来回答。在语料库规模下,为每个文档保留一个全精度向量是主要的索引成本,因此检索系统用几个字节的短码替换每个向量——这一步骤称为量化。标准量化器如乘积量化(PQ)选择最能重构原始向量的码。如果单个码同时服务于多个字节预算,则更有用:当其短前缀各自可直接搜索时,部署可以在不重新编码语料库的情况下设置其效率-质量工作点。但在一个目标下训练所有前缀使得早期比特成为各预算之间的折衷——短码改善而全宽码退化。量化到低比特表示(如二进制码)进一步加剧了冲突。我们引入了马特罗什卡哈希表示(MHR),这是一种两阶段过程,将全宽训练与前缀组织分开。MHR首先学习一个更长的二进制码,然后冻结模型并训练额外的零初始化残差码适配器,用于可直接搜索的前缀。文档以每坐标一位存储,而查询像PQ一样保持连续logits以获得足够的表达能力。我们使用FAISS FastScan实现搜索过程。在MS MARCO上训练并零样本迁移到七个BEIR数据集,MHR在32字节下达到.5561 NDCG@10和.6535 Recall@100,超过了同预算的最佳基线。在较低预算下优势更为明显。相同的码还增强了两种常见流程:为全精度重排序生成候选列表,以及修剪低存储图索引如LEANN。

英文摘要

Retrieval-augmented generation (RAG) depends on dense retrieval: each document is stored as a learned vector, and a query is answered by finding its nearest neighbors in that vector space. Keeping one full-precision vector per document is the dominant index cost at corpus scale, so retrieval systems replace each vector with a short code of a few bytes---a step called quantization. Standard quantizers such as product quantization (PQ) pick the code that reconstructs the original vector most closely. A single code is even more useful if it serves several byte budgets at once: when its short prefixes are each directly searchable, a deployment can set its efficiency--quality operating point without re-encoding the corpus. But training all prefixes under one objective makes the early bits a compromise across budgets---short codes improve while the full-width code degrades. Quantization to low-bit representation, such as binary codes, further sharpens the conflict. We introduce Matryoshka Hash Representations (MHR), a two-stage procedure that separates full-width training from prefix organization. MHR first learns a longer binary code, then freezes the model and trains additional zero-initialized residual code adaptors for directly searchable prefixes. Documents are stored at one bit per coordinate, while queries keep continuous logits like PQ to attain sufficient expressivity. We implement the search process with FAISS FastScan. Trained on MS MARCO and zero-shot transferred to seven BEIR datasets, MHR reaches .5561 NDCG@10 and .6535 Recall@100 at 32 bytes, surpassing the best baseline of the same budget. The advantage is more pronounced in lower budgets. The same code also strengthens two common pipelines: shortlisting candidates for full-precision reranking, and pruning a low-storage graph index such as LEANN.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑