arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式文档检索的标记级信用分配优化

Token-Level Credit Assignment Optimization for Generative Document Retrieval

Xinpeng Zhao, Yang Liu, Ran Chen, Xinyu Ma, Daiting Shi, Pengjie Ren, Zhumin Chen, Zhaochun Ren, Xin Xin

arXiv 2608.12049首次发表:更新:

AI 中文总结

针对生成式检索中标记级生成与文档级监督不匹配问题,提出采用标记级相关性奖励的细粒度强化学习框架,实验显示其优于序列级奖励基线。

AI 中文摘要

生成式检索模型通过自回归地生成文档标识符(DocID)来执行文档检索,该过程自然形成一个序列决策问题,其中每个解码步骤选择一个DocID标记,完整的标记序列决定了被检索的文档。然而,检索有效性通常仅在完整DocID生成后才进行评估,这造成了标记级生成与文档级相关性监督之间的不匹配。因此,现有的生成式检索强化学习方法大多依赖序列级奖励,将相同的文档级反馈传播到所有解码步骤。这种粗粒度的反馈使得难以识别哪些标记决策对检索成功或失败负责。在本研究中,我们提出了一种用于生成式检索的细粒度强化学习框架,该框架采用标记级相关性奖励。我们没有为整个生成的DocID分配单一奖励,而是通过测量每个标记决策如何改变对应生成轨迹的预期检索质量来估计步骤级奖励。这使得信用分配更加精确,并鼓励策略倾向于那些对文档级相关性有更直接贡献的标记决策。我们进一步开发了针对DocID生成过程量身定制的实用奖励估计策略,并将其整合到策略优化框架中。在检索基准上的实验表明,我们的方法始终优于序列级奖励基线,证明了细粒度监督在使自回归DocID生成与检索目标对齐方面的有效性。

英文摘要

Generative retrieval models perform document retrieval by autoregressively generating document identifiers (DocIDs). This process naturally forms a sequential decision problem, i.e., the model makes a sequence of token-level decisions, selecting a DocID token at each decoding step, with the resulting complete sequence identifying the retrieved document. However, relevance feedback is available only after the complete DocID has been generated and mapped to a document, resulting in a granularity mismatch between token-level generation decisions and document-level retrieval supervision. Consequently, existing reinforcement learning methods for generative retrieval rely on sequence-level rewards, assigning the same document-level relevance signal to every decoding step. Such uniform credit assignment obscures the contribution of individual token decisions, making it difficult to identify which decisions contribute to retrieval success or failure. In this paper, we propose Token-Level Credit Assignment for Generative Retrieval (TCA), a fine-grained reinforcement learning framework that aligns the granularity of credit assignment with that of autoregressive DocID generation. Unlike assigning a single reward to an entire generated DocID, TCA derives fine-grained rewards by comparing the hidden-state trajectory of each generated DocID with the gold DocID trajectory obtained from a frozen reference model. These trajectory-based rewards provide differentiated feedback across decoding steps, allowing the policy to reinforce generation paths that remain aligned with the target DocID. Moreover, TCA decouples token-level credit assignment from policy optimization and can be instantiated with both GRPO and PPO. Experiments on benchmarks show that our method consistently outperforms baselines, demonstrating the effectiveness of fine-grained supervision for aligning DocID generation.

Commentsaccepted by CIKM 2026

DOI:10.1145/3799682.3840833

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑