发表机构
Institut de Recherche en Informatique de Toulouse (IRIT)(图卢兹信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PDMR提出段落驱动的多标识生成式检索框架,通过为文档多个段落分配标识符实现多入口表示,并在NQ320K和MS MARCO Document上取得领先性能。
AI 中文摘要
生成式检索(GR)模型将查询直接映射到文档标识符,通过自回归标识符生成取代了对外部稀疏或稠密索引的传统检索。然而,大多数生成式检索框架依赖于单一标识符假设,将每个文档映射到单个目标序列。这迫使模型用一个序列来表示所有文档内容。由于文档通常是多方面的,这可能导致有损表示并降低对查询变化的鲁棒性,其中多个查询意图必须竞争单一的生成访问路径。在这项工作中,我们引入了段落驱动的多标识检索(PDMR),这是一种生成式检索框架,通过多个段落级标识符来表示文档。PDMR对每个文档进行分段,并为每个选定的段落分配一个标识符,这为检索同一文档提供了多个语义入口点。这种多入口表示允许模型将查询与特定的语义方面对齐,从而减少对单个文档级目标的依赖。为了解决这种一对多映射的监督模糊性,我们将训练表述为多目标学习问题,并探索一种旨在将概率质量分配到多个有效段落级标识符上的目标函数。我们在NQ320K和MS MARCO Document上评估了PDMR。在NQ320K上,PDMR在Recall@1和MRR@100上优于强大的生成式和非生成式基线。在MS MARCO Document上,PDMR在所报告的方法中取得了最佳的Recall@1和MRR@10,同时在Recall@10上保持竞争力。受控消融进一步表明,段落级监督、标识符设计、训练查询增强和多目标学习贡献了互补的增益。
英文摘要
Generative Retrieval (GR) models map queries directly to document identifiers, replacing conventional retrieval over external sparse or dense indexes with autoregressive identifier generation. However, most generative retrieval frameworks rely on a single-identifier assumption, mapping each document to a single target sequence. This forces the model to represent all document content with one sequence. Since documents are often multi-faceted, this can lead to lossy representations and reduced robustness to query variation, where multiple query intents must compete for a single generative access path. In this work, we introduce Passage-Driven Multi-ID Retrieval (PDMR), a generative retrieval framework that represents documents through multiple passage-level identifiers. PDMR segments each document and assigns one identifier to each selected passage, which provides multiple semantic entry points for retrieving the same document. This multi-entry representation allows the model to align queries with specific semantic facets, thereby reducing the dependence on a single document-level target. To address the supervision ambiguity of this one-to-many mapping, we formulate training as a multi-target learning problem and explore an objective function designed to distribute probability mass across multiple valid passage-level identifiers. We evaluate PDMR on NQ320K and MS MARCO Document. On NQ320K, PDMR improves over strong generative and non-generative baselines on Recall@1 and MRR@100. On MS MARCO Document, PDMR achieves the best Recall@1 and MRR@10 among the reported methods, while remaining competitive on Recall@10. Controlled ablations further show that passage-level supervision, identifier design, training-query augmentation, and multi-target learning contribute complementary gains.