发表机构
DePaul University; Comcast Technology AI(德保罗大学; 康卡斯特技术AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种服务时路由门控,在协同与LLM生成画像的推荐模型间按需分配用户,在5%NDCG损失预算内提升6.5%的Novelty@10,实现可控新颖性-相关性权衡。
AI 中文摘要
大型语言模型(LLM)能够为推荐系统构建丰富的语义用户画像,但这类画像的生成成本较高,且未必适合在所有场景中统一部署。我们研究是否可以在生产推荐管道中,仅针对特定用户选择性调用LLM生成的画像。利用一个涵盖电影、电视节目和体育内容的真实流媒体数据集,我们训练了一个服务时路由门控,该门控将每个用户分配给协同序列推荐模型或由LLM生成画像驱动的推荐模型。该门控仅使用服务时特征,学习识别那些通过画像路由能够提升Novelty@10且保持排序相关性的用户。路由阈值控制用户被发送到生成模型的激进程度,从而展现出可调节的新颖性-相关性权衡。在整体NDCG损失预算为5%的情况下,学习得到的门控将Novelty@10提升了6.5%,同时路由了12.5%的用户,在相当的相关性成本下优于简单的启发式路由和随机路由策略。这些结果表明,LLM生成的用户画像可以作为协同推荐的受控补充,而与非生成式语义画像的对比结果则表明,收益源于选择性路由而非LLM生成本身。
英文摘要
Large language models (LLMs) enable rich semantic user profiles for recommendation, but such profiles are more expensive to generate and are not necessarily desirable to deploy uniformly. We study whether LLM-generated profiles can instead be invoked selectively within a production recommendation pipeline. Using a real-world streaming dataset covering movies, TV shows, and sports content, we train a serving-time routing gate that assigns each user to either a collaborative sequential recommendation model or a recommendation model driven by an LLM-generated profile. The gate uses only serving-time features and learns to identify users for whom profile-based routing can increase Novelty@10 while preserving ranking relevance. A routing threshold controls how aggressively users are sent to the generative model, exposing a tunable novelty--relevance trade-off. At an overall NDCG-loss budget of 5\%, the learned gate increases Novelty@10 by 6.5\% while routing 12.5\% of users, outperforming simple heuristic and random routing policies at comparable relevance cost. These results show that LLM-generated user profiles can serve as a controllable complement to collaborative recommendation, while results with non-generative semantic profiles indicate that the benefit stems from selective routing rather than LLM generation alone.