arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18796cs.IR

TSGR:淘宝搜索生成式检索

TSGR: Taobao Search Generative Retrieval

Tianyu Zhan, Gui Ling, Tong Xiong, Kunhai Lin, Yang Wang, Kaixuan Zhang, Zhihong Chen, Yuliang Yan, Dan Ou, Shengyu Zhang, Haihong Tang, Bo Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

淘宝搜索生成式检索(TSGR)旨在解决现有生成式检索系统对商品业务价值不敏感的问题。它通过引入查询感知并行SID和价值感知排名模块,将价值感知融入商品表示与候选排名。实验证明该框架有效提升了检索性能及业务指标。

中文摘要 AI 辅助

生成式检索(GR)通过训练单个自回归模型直接生成目标商品的语义ID(SID),在工业电子商务搜索中显示出强大潜力。然而,现有GR系统主要针对语义匹配进行优化,对商品业务价值不敏感:SID构建无价值感知,候选商品排名时未利用商品辅助信息。因此,高价值商品在检索阶段常被遗漏或排名靠后,限制了下游业务影响。在淘宝搜索等工业场景中,这一限制尤为关键,因为业务目标是系统设计的核心。为解决此问题,我们提出了淘宝搜索生成式检索(TSGR),这是一个统一的生成式检索框架,将价值感知纳入商品表示和候选商品排名。对于商品表示,TSGR引入了查询感知并行SID(QP-SID);对于候选商品排名,我们引入了价值感知排名模块(VRM)。离线实验表明,TSGR在HR@1000上提高了9.16%,在线A/B测试进一步验证了其有效性,在IPV、交易数量和GMV方面分别提高了0.43%、1.12%和1.64%。

英文摘要

Generative retrieval (GR) has demonstrated strong promise for industrial e-commerce search by training a single autoregressive model to directly generate the Semantic IDs (SIDs) of target items. However, existing GR systems are primarily optimized for semantic matching and remain insensitive to item business value: SID construction is value-unaware, and candidates are ranked without access to item side-info. Consequently, high-value items are often missed or deprioritized at the retrieval stage, limiting downstream business impact. This limitation is particularly critical in industrial settings such as Taobao Search, where business objectives are central to system design. To address this, we propose $\textbf{T}$aobao $\textbf{S}$earch $\textbf{G}$enerative $\textbf{R}$etrieval ($\textbf{TSGR}$), a unified generative retrieval framework that incorporates value awareness into both item representation and candidate ranking. 1) For item representation, TSGR introduces $\textbf{Query-aware Parallel SID (QP-SID)}$, which encodes query-conditioned value orderings into the SID construction by building parallel codebooks derived from query-item statistics, so that higher-value and query-relevant items are assigned better token indices. 2) For candidate ranking, we introduce a $\textbf{Value-aware Ranking Module (VRM)}$ that is built upon and jointly optimized with the GR, enabling a single model to seamlessly serve as both retriever and pre-ranker without a dedicated pre-ranking stage. A progressive training pipeline further aligns the model with semantic relevance, user preferences, and business objectives. Offline experiments show that TSGR achieves a 9.16\% improvement in HR@1000, and online A/B tests further validate its effectiveness, yielding gains of +0.43\% in IPV, +1.12\% in Transaction Count, and +1.64\% in GMV. TSGR has been fully deployed in production.

↑