解耦生成式检索中的范式、标识符与解码
Disentangling Paradigm, Identifier, and Decoding in Generative Retrieval
浏览论文内容
中文总结 AI 辅助
本研究在固定标识符长度和训练预算下,解耦比较生成式检索中的范式、标识符与解码,发现解码方式显著影响扩散模型性能,且范式比较需按各自最优方案进行。
中文摘要 AI 辅助
生成式检索训练语言模型生成相关文档的标识符。近期工作用扩散模型替代自回归解码器,但同时改变了标识符、训练方案和解码方式,因此差异不能归因于范式本身。在NQ320K和MS300K数据集上,我们使用残差量化、乘积量化和随机标识符训练了自回归、掩码扩散和块扩散模型。在固定标识符长度和训练预算的条件下,我们以多种方式解码每个模型。仅改变解码方式就能使扩散模型的Hit@1提升6.6至13.7个百分点。我们提出的参考扩散解码方法——生成后匹配(generate-and-match)——首先生成标识符,然后检索最接近的语料库标识符。该生成的标识符对NQ320K中14-21%的查询是正确的。我们测试了单遍评分(one-pass scoring)来解码扩散检索器:模型一次读取完全掩码的标识符,每个文档根据其编码的概率进行评分。在12种设置中的11种里,该方法匹配或超越了生成后匹配。自回归模型在Hit@1上仍然领先;在NQ320K上,这一领先来自模型本身,而非束搜索。从单个采样标识符出发,单遍评分消除了掩码扩散相对于束搜索差距的46-83%;从生成后匹配出发,最多消除四分之一。在NQ320K上,每种范式在很大程度上记住了哪个标识符回答哪个查询:随机标识符保留了残差量化标识符Hit@1的83-90%。在那里,乘积量化标识符在自回归模型中比残差量化标识符领先3.4个百分点,在扩散模型中领先-0.7至+3.6个百分点;跨解码方式,自回归的差距比扩散大1.5-2.3个百分点,接近我们2个百分点的阈值。范式比较必须报告每种范式在其自身方案和最佳解码下的表现。
英文摘要
Generative retrieval trains a language model to generate the identifier of a relevant document. Recent work replaces the autoregressive decoder with diffusion, but changes identifiers, training recipe and decoding at once, so differences cannot be credited to the paradigm. On NQ320K and MS300K, we train autoregressive, masked-diffusion and block-diffusion models with residual-quantised, product-quantised and random identifiers. With identifier length and training budget fixed, we decode each model in several ways. Decoding alone moves a diffusion model's Hit@1 by 6.6 to 13.7 points. Our reference diffusion decoding, generate-and-match, generates an identifier, then retrieves the closest corpus identifiers. The generated identifier is right for 14-21% of NQ320K queries. We test one-pass scoring to decode diffusion retrievers: the model reads a fully masked identifier once, and each document is scored by its codes' probabilities. It matches or beats generate-and-match in 11 of 12 settings. Autoregressive models still lead in Hit@1; on NQ320K, the lead comes from the model, not beam search. Starting from one sampled identifier, one-pass scoring removes 46-83% of masked diffusion's deficit to beam search; from generate-and-match, at most a quarter. On NQ320K, every paradigm largely memorises which identifier answers which query: random identifiers keep 83-90% of the Hit@1 of residual-quantised ones. There, product-quantised identifiers lead residual-quantised ones by 3.4 points in the autoregressive model and by -0.7 to +3.6 in diffusion models; across decodings, AR's gap exceeds diffusion's by 1.5-2.3 points, around our 2-point threshold. Paradigm comparisons must report each paradigm at its own recipe and best decoding.
发表机构
- Artefact Research Center(人工制品研究中心)
机构由 AI 辅助整理,请以论文原文为准。