AI 中文总结
研究针对大规模运行RecGPT的挑战,提出RecGPT-V3这一有状态混合模态推荐系统。通过内存中心、混合模态基础模型和潜在意图推理等方法,在淘宝“猜你喜欢”推荐中取得多指标提升及资源消耗降低的成果。
AI 中文摘要
大型语言模型正在改变推荐系统,从匹配历史行为中的共现模式转向推理行为背后的意图。RecGPT-V1以用户理解为核心开创了这一范式,RecGPT-V2通过协同多智能体推理进行扩展,二者都已投入生产并带来用户体验和商业成果的提升。然而,大规模运行RecGPT存在三个挑战:无状态行为建模、标签到项目的信息瓶颈以及低效的显式推理。我们提出了RecGPT-V3,一种有状态的混合模态推荐系统,它通过自然语言推理开放世界知识,通过语义ID进行具体项目定位。一个内存中心维护结构化、不断演变的用户内存,将长期行为提炼为压缩单元,将用户建模计算减少55.8%。混合模态基础模型允许大语言模型联合推理文本标签和语义ID,开辟进入项目空间的高带宽通道。潜在意图推理将冗长的推理过程内化为紧凑的可学习潜在令牌,可解码为可读解释,将输出令牌成本降低200倍。在淘宝的“猜你喜欢”推荐中部署后,RecGPT-V3在大规模在线A/B测试中取得了持续的收益:商品详情页浏览量提升1.28%,点击率提升1.00%,转化率提升1.97%,商品交易总额提升3.97%,同时将端到端服务资源消耗降低52.4%。
英文摘要
Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior modeling, where each request reprocesses full user history, wasting computation and discarding prior analysis; (2) a tag-to-item information bottleneck, where natural-language tags form a lossy channel between user understanding and item grounding; and (3) inefficient explicit reasoning, whose lengthy chain-of-thought incurs untenable latency and compute overhead. We present RecGPT-V3, a stateful, hybrid-modal recommender that reasons over natural language for open-world knowledge and Semantic IDs (SIDs) for concrete item grounding. A Memory Hub maintains structured, continually evolving user memory that distills long-horizon behavior into condensed units, cutting user-modeling computation by 55.8%. A Hybrid-modal Foundation Model allows the LLM jointly reason over text tags and SIDs, opening a high-bandwidth channel into the item space. Latent Intent Reasoning internalizes verbose rationales into compact learnable latent tokens that remain decodable into readable explanations, lowering output token cost by 200x. Deployed in Taobao's "Guess What You Like" feed, RecGPT-V3 achieves consistent gains in large-scale online A/B tests: IPV +1.28%, CTR +1.00%, TC +1.97%, GMV +3.97%, while cutting end-to-end serving resource consumption by 52.4%.
CommentsTechnique Report