arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

思考以个性化:统一推理与检索的以用户为中心的个性化密集检索

Think-to-Personalize: Unifying Reasoning and Retrieval for User-Centric Personalized Dense Retrieval

Angqing Jiang, Gaoming Zhang, Jianchun Song, Kena Qi, Dayao Chen, Wei Lin, Defu Lian

arXiv 2608.18855首次发表:更新:

AI 中文总结

本文提出TTP框架,将用户意图推理与密集检索统一,经两阶段训练后在电商场景实验及在线测试中均取得优于基线的效果,提升了订单量。

AI 中文摘要

密集检索已成为现代本地生活电商搜索的核心,其将查询与商品编码成语义嵌入空间。尽管近期进展已从基于BERT的嵌入模型过渡到大型语言模型(LLMs),但多数方法仍将LLMs视为静态文本编码器,忽略其固有的推理能力。此外,标准密集检索模型仍以查询为中心,在电商场景中,稀疏且模糊的查询会产生意图差距,仅能通过丰富的用户历史上下文弥合;而现有个性化检索方法通常依赖隐式嵌入交互,缺乏从噪声历史行为中有效消歧用户意图的推理能力。为应对这些挑战,本文提出Think-to-Personalize(TTP),这一新颖框架将显式以用户为中心的意图推理与密集检索相统一。通过对用户历史购买序列进行推理,TTP显式推导潜在的个性化需求并生成意图增强的查询,随后将其编码为统一的密集嵌入。具体而言,本文设计了两阶段训练范式:(1)监督微调(SFT)阶段,建立冷启动能力;(2)强化学习(RL)阶段,使用分组相对策略优化(GRPO)使推理过程与检索效用对齐。在自有数据集及公开基准上的大量实验表明,TTP显著优于当前最优基线;此外,在线A/B测试中,其订单量提升了0.46%,验证了其实际有效性,为推理驱动的个性化密集检索建立了新范式。

英文摘要

Dense retrieval has become a cornerstone of modern local-lifestyle e-commerce search by encoding queries and items into semantic embedding spaces. While recent advancements have transitioned from BERT-based embedding models to Large Language Models (LLMs), most approaches still treat LLMs as static text encoders, neglecting their inherent reasoning capabilities. Furthermore, standard dense retrieval models remain query-centric, which is insufficient in e-commerce scenarios where sparse and ambiguous queries create an intent gap that can only be bridged by the rich context of user history. Meanwhile, existing personalized retrieval methods typically rely on implicit embedding interactions, which lack the reasoning capability to effectively disambiguate user intent from noisy historical behaviors. To address these challenges, we propose Think-to-Personalize (TTP), a novel framework that unifies explicit user-centric intent reasoning with dense retrieval. By reasoning over the user's historical purchase sequence, TTP explicitly deduces latent personalized needs and generates an intent-enhanced query, which is then encoded into a unified dense embedding. Specifically, we design a two-stage training paradigm: (1) a Supervised Fine-Tuning (SFT) stage that establishes cold-start capabilities; and (2) a Reinforcement Learning (RL) stage that aligns the reasoning process with retrieval utility using Group Relative Policy Optimization (GRPO). Extensive experiments on both proprietary and public benchmarks demonstrate that TTP significantly outperforms state-of-the-art baselines. Furthermore, in online A/B tests, it achieved a +0.46% lift in order volume, validating its practical effectiveness and establishing a new paradigm for reasoning-driven personalized dense retrieval.

CommentsAccepted at CIKM 2026. 11 pages, 8 figures, and 9 tables

DOI:10.1145/3799682.3840844

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑