arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式信息检索中语义ID空间的系统研究

A Systematic Study of Semantic ID Spaces for Generative Information Retrieval

Alexia Allal, Hicham Randrianarivo, Sylvain Lamprier

arXiv 2610.08732首次发表:更新:

发表机构

Artefact Research Center; Angers University(人工制品研究中心; 昂热大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统探究生成式信息检索中DocID的语义设计,提出统一量化框架及无需训练的固有指标,并在MS MARCO 300K和NQ320K上验证结构属性对检索效果的影响。

AI 中文摘要

生成式信息检索(GIR)已成为一种变革性范式,将文档检索从传统的“检索-排序”流程转变为序列到序列生成,其中模型直接预测文档标识符(DocID)。虽然这些DocID的语义设计已知对性能至关重要,但一个基本问题仍未得到充分探索:什么构成了一个好的DocID?当前方法严重依赖计算成本高昂的下游评估,阻碍了系统分析和快速迭代。在本工作中,我们通过提出一项关于定义有效数值DocID的属性、指标和权衡的综合研究来解决这一挑战。具体而言,我们的贡献有三方面:首先,我们提出了一个统一框架,将乘积量化(PQ)和残差量化(RQ)及其混合变体统一在单一设计空间中。这使我们能够系统地研究关键DocID属性,如层次性与并行性,以及DocID长度和码本大小等超参数的影响。其次,我们定义了一套无需训练的固有指标,用于量化DocID质量并评估结构保真度,而无需完整模型训练的开销。通过在MS MARCO 300K和NQ320K上的广泛实验,我们分析了这些结构属性如何影响检索效果。

英文摘要

Generative Information Retrieval (GIR) has emerged as a transformative paradigm, shifting document retrieval from a traditional "retrieve-and-rank" workflow to sequence-to-sequence generation, where a model directly predicts document identifiers (DocIDs). While the semantic design of these DocIDs is known to be critical for performance, a fundamental question remains under-explored: what makes a good DocID? Current approaches rely heavily on computationally expensive downstream evaluations, hindering systematic analysis and rapid iteration. In this work, we address this challenge by presenting a comprehensive study on the properties, metrics, and trade-offs that define effective numerical DocIDs. Specifically, our contributions are threefold: First, we propose a unified framework that unifies Product Quantization (PQ) and Residual Quantization (RQ), and their hybrid variants within a single design space. This enables us to systematically study key DocID properties, such as hierarchy versus parallelism, as well as the impact of hyperparameters like DocID length and codebook size. Second, we define a suite of training-free, intrinsic metrics, to quantify DocID quality and evaluate structural fidelity without the overhead of full model training. Through extensive experiments on MS MARCO 300K and NQ320K, we analyze how these structural properties influence retrieval effectiveness.

Comments8 pages, 3 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑