arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19595cs.IR

SSR-GRPO:将监督与语义标识符集成到强化学习中用于电商领域的密集检索

SSR-GRPO: Integrating Supervision and Semantic IDs into Reinforcement Learning for Dense Retrieval in E-commerce

Guangxin Song, Xing Fang, Mingmin Jin, Jing Wang, Bokang Wang, Zhentao Song, Junjie Bai, Jianbo Zhu

AI总结:

该研究针对电商嵌入检索的语义处理难题,提出SSR-GRPO模型,通过双视角相关性评估、难负样本挖掘等策略优化R-GRPO,经实验验证有效并已部署于大规模电商平台。

AI中文摘要:

基于嵌入的检索(EBR)在电商搜索中至关重要,但往往难以处理复杂语义。尽管近期方法常对大型语言模型(LLM)进行微调以用于表示学习,但它们通常缺乏处理复杂和隐式语义的稳健机制。最近的Retrieval-GRPO(R-GRPO)虽将强化学习引入密集检索,却因批量采样有限导致Top-K候选存在噪声,且因使用相似训练的LLM作为奖励模型产生有偏相关性评估。为解决这些问题,我们提出带有语义标识符(SIDs)的监督检索-GRPO(SSR-GRPO)。具体而言,我们的方法首先提出双视角相关性评估框架,利用量化学习生成的语义标识符(SIDs)与密集表示向量生成更无偏的相关性得分;此外,利用生成SIDs的层次相似关系,挖掘出一组难负样本,其有两个用途:(1)设计集成到R-GRPO中的掩码函数,有效过滤组内噪声样本;(2)构建由正负样本对组成的Retrieval-DPO任务,使模型能从成对视角捕捉细粒度语义差异。通过整合这些优化策略,我们提出SSR-GRPO。大量离线和在线实验验证了SSR-GRPO的有效性,且该模型已部署在一个大规模电商平台上。

英文摘要:

Embedding-based retrieval (EBR) is pivotal in e-commerce search but often struggles with complex semantics. While recent methods often fine-tune large language models (LLMs) for representation learning, they typically lack robust mechanisms for handling complex and implicit semantics. While Retrieval-GRPO (R-GRPO) recently introduced reinforcement learning to dense retrieval, it suffers from noisy top-K candidates due to limited batch sampling and biased relevance assessments caused by using similarly trained LLMs as reward models. To tackle these issues, we propose Supervised Retrieval-GRPO with Semantic Identifiers (SSR-GRPO). Specifically, our method first proposes a dual-perspective framework for relevance assessment. It leverages both Semantic Identifiers (SIDs) produced by quantization learning and dense representation vectors to generate more unbiased relevance scores. Furthermore, leveraging the hierarchical similarity relationships of the generated SIDs, we mine a set of hard negative samples that serve two purposes: (1) to design a masking function integrated into R-GRPO, effectively filtering intra-group noisy samples; and (2) to construct a Retrieval-DPO task composed of positive and negative sample pairs, enabling the model to capture fine-grained semantic distinctions from a pair-wise perspective. By integrating these optimization strategies, we propose SSR-GRPO. Extensive offline and online experiments validate SSR-GRPO's effectiveness, and it has been deployed on a large-scale e-commerce platform.

↑