arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39312cs.IR

学习多分辨率相关性用于分层生成式检索

Learning Multiresolution Relevance for Hierarchical Generative Retrieval

  • Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
  • Meituan(美团)

机构由 AI 辅助整理,请以论文原文为准。

Weihao Shen, Wei Chen, Fuwei Zhang, Guojun Liu, Qingsong Hua, Wei Lin, Fuzhen Zhuang

AI总结:

提出RARS方法,通过将文档级相关性分解为层级条件分布来监督查询编码器,在自回归生成式检索中提升多语言ESCI数据集上的检索性能,且推理时无需额外开销。

AI中文摘要:

基于语义标识符(SIDs)的生成式检索在文档层次结构上做出连续决策。对于同一查询,相关文档可能共享粗粒度前缀,并在更细粒度层级上产生分歧,且分支模式因查询而异。这些路径揭示了相关性如何在连续的细化过程中分布,然而标准的全SID监督将这些路径视为独立的训练目标。为使这种分配显式化,我们将多分辨率相关性表述为由单一文档级相关性度量在SID层级上诱导的一致条件分布。我们提出RARS(分辨率对齐的相关性监督),利用由此产生的细化层级分布来监督共享的查询表示。RARS对前缀上的文档相关性进行聚合,并训练一个前缀条件预测器来在兄弟分支之间分配相关性。所有承载相关性的子节点参与局部竞争,每个局部损失按其到达父节点的相关性质量加权。该目标训练查询编码器同时捕获相关文档共享的粗粒度结构及其更细的分支分配。训练后丢弃预测器,在推理时保留标准自回归检索。在三个多语言ESCI区域上的实验表明,在自回归解码下,与匹配的全SID训练相比,RARS取得了一致的改进。在常见检索规则下,RARS也优于分组软目标、解码器软目标和采样树监督。这些改进在替代标识符结构和相关性定义下依然保持。代码可在以下网址获取:this https URL

英文摘要:

Generative retrieval with semantic identifiers (SIDs) makes successive decisions over a document hierarchy. Relevant documents for the same query may share coarse prefixes and diverge at finer depths, with branching patterns varying across queries. These paths reveal how relevance is distributed across successive refinements, yet standard full-SID supervision treats them as separate training targets. To make this allocation explicit, we formulate multiresolution relevance as consistent conditional distributions induced by a single document-level relevance measure across the SID hierarchy. We introduce \textbf{RARS}, \textbf{R}esolution-\textbf{A}ligned \textbf{R}elevance \textbf{S}upervision, which uses the resulting refinement-level distributions to supervise a shared query representation. RARS aggregates document relevance over prefixes and trains a prefix-conditioned predictor to allocate relevance among sibling branches. All relevance-bearing children participate in local competition, and each local loss is weighted by the relevance mass reaching its parent. This objective trains the query encoder to capture both the coarse structure shared by relevant documents and their finer branch allocations. The predictor is discarded after training, preserving standard autoregressive retrieval at inference. Experiments on three multilingual ESCI locales show consistent improvements over matched full-SID training under autoregressive decoding. RARS also outperforms grouped soft-target, decoder soft-target, and sampled-tree supervision under a common retrieval rule. The gains persist across alternative identifier structures and relevance definitions. Code is available at: https://github.com/Nevaeh7/RARS

↑