AI 中文总结
该研究针对将预训练LLM转化为生成式检索器的挑战,设计了SnapLGR系统,通过三项核心优化实现了Snapchat短视频推荐效果的提升。
AI 中文摘要
预训练大语言模型(LLM)是极具潜力的检索引擎,因为它们兼具丰富的语义先验、强大的序列建模能力以及良好的缩放特性。然而,将预训练LLM转化为生产部署中的生成式检索器会带来诸多挑战:模型必须学习预训练阶段未出现的内部物品词汇表,且需在严格的延迟和成本约束下生成有效的物品标识符。我们通过设计并推出SnapLGR解决了这些挑战,这是一个用于Snapchat短视频推荐的基于LLM的生成式检索系统。该系统围绕三项核心设计构建:第一,我们从多模态物品嵌入中构建语义标识符(SIDs),并通过基于个性化PageRank(PPR)的共同参与对比学习对其进行增强,从而提升码本利用率、减少冲突并注入协同信号;第二,我们使用持续预训练(CPT)将引入的SID标记进行基础化处理,之后再对用户交互序列进行监督微调(SFT);第三,我们通过基于TensorRT-LLM CUDA的集束搜索以及去中心化工作者循环架构,使SnapLGR的服务具备实用性。在在线A/B测试中,与现有的TIGER式生成式检索基线相比,推出的系统使观看时长提升了0.37%,花费时间提升了0.09%,深度会话提升了0.18%,深度会话独立用户提升了0.11%。随后,我们在固定分词器下对这种离线差距进行分解,并量化了模型架构、缩放和预训练带来的增益。总体而言,我们的部署表明,成功落地SnapLGR需要在表示学习、词汇基础化以及高效训练与服务之间进行协同设计。
英文摘要
Pretrained large language models (LLMs) are promising retrieval engines because they combine rich semantic priors, strong sequence modeling capabilities, and favorable scaling behavior. However, turning a pretrained LLM into a generative retriever in production deployment raises several challenges: the model must learn an internal item vocabulary that was absent from pretraining, and generate valid item identifiers under strict latency and cost constraints. We address these challenges through the design and launch of SnapLGR, an LLM-based generative retrieval system for short-video recommendation at Snapchat. The system is built around three main designs. First, we construct semantic identifiers (SIDs) from multimodal item embeddings and enhance them with Personalized PageRank (PPR)-based co-engagement contrastive learning, resulting in improved codebook utilization, reduced collisions, and infused collaborative signal. Second, we use continued pretraining (CPT) to ground the introduced SID tokens before supervised fine-tuning (SFT) on user interaction sequences. Third, we make SnapLGR serving practical through TensorRT-LLM CUDA-backed beam search and a decentralized worker-loop architecture. In a live A/B test, the launched system increased View Time by 0.37%, Time Spent by 0.09%, Deep Sessions by 0.18%, and Deep Sessions Unique User by 0.11% relative to the existing TIGER-style generative retrieval baseline. We then decompose this offline gap under a fixed tokenizer and quantify the gains due to model architecture, scaling, and pretraining. Overall, our deployment shows that successful production SnapLGR requires joint design across representation learning, vocabulary grounding, and efficient training and serving.