arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从检索内容中学习:面向语义检索的在线强化学习微调

Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval

Shaowei Wei, Chong Huang, Songtao Fang, Jin Zhang, Zhuojun Wang, Chengfu Huo

arXiv 2608.30753首次发表:更新:

发表机构

Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对电商检索中双编码器与重排器的目标不匹配问题,提出PAO选择性RL优化方法,在冻结文档索引下提升检索性能且优于标准RL与蒸馏基线。

AI 中文摘要

在大规模电商检索场景中,双编码器检索器针对对比相似度进行优化,而下游重排器则捕捉更细粒度的相关性偏好;这种目标不匹配限制了端到端检索质量。强化学习提供了利用奖励模型反馈适配检索器的途径,但我们发现标准策略梯度更新会损害嵌入几何结构,尤其是在工业场景下文档索引必须保持冻结时。为解决该问题,我们提出PAO(仅正优势)这一选择性强化学习优化方法。我们的分析表明,在冻结的高维空间中不加区分地惩罚负样本(推离)会破坏预训练的语义流形。PAO仅对具有正优势的检索项应用梯度更新,有效将查询嵌入拉向高奖励区域,同时保留全局拓扑稳定性。在大规模工业数据集和公开基准上的实验表明,PAO显著优于标准强化学习和蒸馏基线方法。

英文摘要

In large-scale e-commerce retrieval, dual-encoder retrievers are op- timized for contrastive similarity, whereas downstream rerankers capture finer-grained relevance preferences; this objective mis- match limits end-to-end retrieval quality. Reinforcement Learning offers a way to use reward-model feedback for retriever adaptation, but we observe that standard policy-gradient updates can degrade embedding geometry, especially when the document index must remain frozen due to industrial constraints. To address this, we propose PAO (Positive-Advantage-Only), a selective RL optimization method. Our analysis reveals that in- discriminate penalization of negative samples (pushing away) in a frozen high-dimensional space disrupts pre-trained semantic man- ifolds. PAO selectively applies gradient updates only to retrieved items with positive advantages, effectively pulling query embed- dings toward high-reward regions while preserving global topo- logical stability. Experiments on both a massive industrial dataset and public benchmarks demonstrate that PAO significantly outper- forms standard RL and distillation baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑