发表机构
Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RELER框架,利用强化学习在嵌入空间中直接优化检索指标,通过vMF采样和条件均值投影减少噪声,在BRIGHT基准及RAG下游任务中显著提升性能。
AI 中文摘要
密集检索模型通常使用对比目标进行训练,这些目标能学习有效的表示,但不会直接优化检索指标或下游任务性能。为解决这一问题,我们提出了RELER(用于检索的强化学习),这是一个强化学习框架,使现有的嵌入模型能够直接在嵌入空间中学习检索,并与特定任务的奖励对齐。我们通过从以归一化编码器输出为中心的von Mises-Fisher(vMF)分布中采样单位长度的查询和文档嵌入动作来训练RELER,将得到的检索或下游结果作为奖励进行评分,并使用带有留一法基线(RLOO)的REINFORCE算法更新编码器。由于高维嵌入空间中的探索容易受到采样噪声的影响,我们进一步提出了条件均值投影(CMP),它将每个采样的嵌入投影到由其编码器输出和与之比较的候选嵌入所张成的低维子空间上,在保持策略梯度期望的同时减少噪声。我们在BRIGHT上评估了RELER,这是一个包含推理密集型查询的基准,对现有嵌入模型仍具挑战性。在对BGE-M3和Qwen3-Embedding骨干进行后训练时,RELER在平均nDCG@10上始终优于InfoNCE和LambdaLoss。我们进一步通过检索增强生成(RAG)评估下游效用,其中我们仅调整查询编码器,同时保持文档索引和生成器固定。在七个QA数据集上,联合优化检索和答案奖励可提高RAG中的平均检索性能和答案质量。
英文摘要
Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.