发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对电商生成式检索的级联训练误差累积与交互建模不足问题,提出联合训练嵌入模型和码本并引入产品聚类监督的方法,有效提升了检索性能。
AI 中文摘要
随着大型语言模型(LLMs)的发展,生成式检索在电商场景中愈发重要。当前主流方法通常采用两阶段训练策略:首先训练产品嵌入模型,再学习将嵌入映射到产品ID的码本。这种级联方法存在两个主要问题:(1)误差累积——若第一阶段的嵌入模型生成有偏差的表示,第二阶段的码本无法纠正这些误差,会降低最终检索性能;(2)码本学习仅依赖产品嵌入,缺乏对查询-产品及产品-产品交互的建模,导致同一聚类的产品可能被码本分配不一致的ID,进一步损害检索准确率。为解决这些问题,本文提出一种联合训练嵌入模型与码本的新方法,并将同一产品聚类信息作为额外监督信号。实验结果表明,该方法可显著提升电商检索性能,同时增强嵌入与码本学习。
英文摘要
With the development of large language models (LLMs), generative retrieval is becoming increasingly important in e-commerce scenarios. Current mainstream approaches typically use a two-stage training strategy: first train a product embedding model, and then learn a codebook that maps embeddings to product IDs. This cascaded approach suffers from two major issues: (1) error accumulation-if the embedding model in the first stage produces biased representations, the codebook in the second stage cannot correct these errors, degrading final retrieval performance; and (2) codebook learning relies solely on product embeddings and lacks modeling of query-to-product and product-to-product interactions. As a result, products belonging to the same cluster may be assigned inconsistent IDs by the codebook, further hurting retrieval accuracy. To address these problems, we propose a novel method that jointly trains the embedding model and the codebook, and incorporates same product cluster information as an additional supervision signal. Experimental results demonstrate that our method significantly improves e-commerce retrieval performance while simultaneously enhancing both embedding and codebook learning.