arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00638cs.IRcs.CL

匹配需双方协作:通过强化学习协同进化生成式检索器

It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning

Runpeng Dai, Kaili Huang, Changsung Kang, Ciya Liao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出CoGR框架,通过两阶段训练与GRPO强化学习协同优化查询、项侧生成器,在多基准数据集上较最强基线F1值提升10.9%与36.1%,实现更优检索效果。

中文摘要 AI 辅助

检索是现代搜索与广告系统的首要环节,需从海量候选项库中筛选出候选集供下游排序与拍卖使用。近期研究日益利用大语言模型(LLMs)通过查询扩展、数据合成及检索反馈训练来改进检索效果,但生成组件通常仅用于查询侧增强,最终匹配仍依赖下游检索器。本文提出CoGR(协同生成式检索框架),训练LLMs直接构建查询侧与项侧的检索表示,每个生成器生成一组紧凑关键词,通过倒排索引直接匹配,可兼容现有基于关键词的检索基础设施。CoGR采用两阶段训练流程:先通过监督微调建立对齐的关键词空间,再通过协同进化强化学习交替优化查询侧与项侧生成器,其中GRPO算法会固定对方索引进行优化。两侧均以相同的查询-项检索F1值为目标:查询侧直接接收检索F1值,项侧则接收反事实边际奖励,该奖励衡量其生成关键词导致的查询侧F1值变化。在10种代表性稀疏、密集及生成式基准模型中,CoGR在内部APP市场数据集与公开WANDS基准上均取得最佳性能,相较于最强基线模型,F1值分别提升10.9%与36.1%。进一步分析显示,训练过程中协同进化稳定,查询与项的关键词空间对齐度逐步提升。

英文摘要

Retrieval is the first stage of modern search and advertising systems, selecting a candidate set from a large item universe for downstream ranking and auction. Recent work increasingly leverages LLMs to improve retrieval through query expansion, data synthesis, and retrieval-feedback training. However, the generative component is typically used for query-side augmentation, while final matching is still delegated to a downstream retriever. We introduce CoGR, a retrieval framework that instead trains LLMs to directly construct retrieval representations on both query and item sides. Each generator produces a compact set of keywords, which are matched directly through an inverted index, preserving compatibility with existing keyword-based retrieval infrastructure. CoGR uses a two-stage training pipeline. Supervised fine-tuning first establishes an aligned keyword space, after which co-evolving reinforcement learning alternately optimizes the query- and item-side generators with GRPO against the opposite side's frozen index. Both sides optimize the same query-to-item retrieval $F_1$ objective: the query side receives retrieval $F_1$ directly, while the item side receives a counterfactual marginal reward measuring the change in query-side $F_1$ caused by its generated keywords. Across 10 representative sparse, dense, and generative baselines, CoGR achieves the best performance on both an internal APP Marketplace dataset and the public WANDS benchmark, improving $F_1$ over the strongest baseline by $10.9\%$ and $36.1\%$, respectively. Further analysis shows stable co-evolution and increasingly aligned query--item keyword spaces over training.

发表机构

  • University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
  • Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

↑