AI 中文总结
研究针对增量视频搜索中意图不明的问题,提出结合语义与协作信号的个性化系统。通过学习文本和ID嵌入空间构建用户表示,注入相似度到排名器。实验表明离线提升了NDCG@10和MRR,在线提高了点击率、转化率和排名位置,还分析了嵌入权衡与质量。
AI 中文摘要
增量视频搜索要求在每次按键后都有高质量的排名,此时意图往往不明确(例如1-3个字符的前缀)。我们提出了一个用于苹果电视搜索的个性化系统,该系统在排名时结合了互补的语义和协作信号。我们的方法学习两个项目嵌入空间:(i)一个基于文本的多语言编码器(TextEmb),通过对比学习在共同参与三元组上进行微调;(ii)一个基于ID的协作嵌入模型(IdEmb),在交互衍生的正样本上进行训练。在服务时,我们根据最近的观看历史构建用户表示,并将基于文本和ID的用户-项目余弦相似度注入到成对的XGBoost排名器中。我们使用时间上留出的离线数据集和为期三周的在线对照实验进行评估。离线时,对于有用户历史的会话,个性化排名器相对于非个性化基线,NDCG@10提高了2.99%,MRR提高了3.30%。切片分析表明,在意图仍在形成的增量搜索中,个性化最为需要:在模糊前缀查询(1-3个字符)上,NDCG@10提升为+8.63%,而在更长、完全指定的查询上为+1.46%。历史更长的用户受益更多:NDCG提升从有1-5个历史项目的用户的+2.13%上升到有51-100个历史项目的用户的+4.37%,尽管这些群体的基线相关性较低(NDCG@10从0.733下降到0.680),这表明在默认排名表现不佳的地方,个性化增加的价值最大。在线时,处理在点击率上有统计学上显著的+1.14%的提升,转化率有+1.23%的提升,转化项目的排名位置有2.91%的改善。我们还通过隔离每个信号来分析语义和协作嵌入之间的覆盖-精度权衡,并在一个有LLM判断相似度标签的留出语料库上评估嵌入质量,以减少点击/曝光偏差。
英文摘要
Incremental video search requires high-quality ranking after each keystroke, where intent is often underspecified (e.g., 1-3 character prefixes). We present a personalization system for Apple TV search that combines complementary semantic and collaborative signals at ranking time. Our approach learns two item embedding spaces: (i) a text-based multilingual encoder (TextEmb) fine-tuned on co-engagement triplets via contrastive learning, and (ii) an ID-based collaborative embedding model (IdEmb) trained on interaction-derived positives. At serving time, we construct user representations from recent watch history and inject text- and ID-based user-item cosine similarities into a pairwise XGBoost ranker. We evaluate with temporally held-out offline datasets and a three-week online controlled experiment. Offline, for sessions with user history, the personalized ranker improves NDCG@10 by 2.99% and MRR by 3.30% over the non-personalized baseline. Slice analyses show that personalization is most needed in incremental search, where intent is still forming: on ambiguous prefix queries (1-3 characters), NDCG@10 lift is +8.63%, versus +1.46% on longer, fully specified queries. Longer-history users benefit more: NDCG lift rises from +2.13% for users with 1-5 history items to +4.37% for users with 51-100, even though baseline relevance is lower for these cohorts (NDCG@10 drops from 0.733 to 0.680), indicating that personalization adds the most value where default ranking underperforms. Online, treatment yields statistically significant gains of +1.14% tap-through rate and +1.23% conversion rate, with a 2.91% improvement in converted-item rank position. We further analyze coverage-precision trade-offs between semantic and collaborative embeddings via ablations isolating each signal, and evaluate embedding quality on a held-out corpus with LLM-judged similarity labels to reduce click/exposure bias.
CommentsAccepted to the Industry Track of the 20th ACM Conference on Recommender Systems (RecSys 2026)