基于LLM的用户画像在生成式流推荐中的价值时机
When LLM-Based User Profiling Adds Value in Production Streaming Recommendation
- DePaul University(德保罗大学)
- Comcast Technology AI(康卡斯特技术AI)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文系统比较了基于LLM与聚合方法在流推荐中构建语义用户画像的四种策略,评估其在不同用户行为和时间窗口下的价值,以确定LLM额外成本的合理时机。
AI中文摘要:
个性化推荐的关键在于如何从历史行为中构建用户表示。在基于内容的推荐中,出现了两种构建语义用户画像的范式。第一种是聚合方法,将用户表示推导为语义项目嵌入的数值聚合。第二种是基于LLM的方法,生成用户偏好的自然语言摘要,并通过文本编码器进行编码。每种范式都可以与近期与历史行为的时间解耦相结合。基于LLM的画像生成比聚合方法昂贵得多,这引出了额外成本何时合理的问题。我们对四种语义用户画像策略进行了系统比较,这些策略在表示类型和时间处理上进行了因子交叉,并在真实生产数据集上进行了评估。比较揭示了这些策略在不同用户行为类型、推荐质量的准确性和超准确性维度以及控制解耦的时间窗口设置上的差异。
英文摘要:
Personalized recommendation depends critically on how user representations are constructed from historical behavior. Two paradigms have emerged for constructing semantic user profiles in content-based recommendation. First, aggregate methods derive user representations as numerical aggregates of semantic item embeddings. Second, LLM-based methods generate natural-language summaries of user preferences and encode them through a text encoder. Each paradigm can be combined with temporal disentanglement of recent versus historical behavior. LLM-based profile generation is significantly more expensive than aggregate approaches, raising the question of when this additional cost is justified. We present a systematic comparison of four semantic user-profiling strategies, factorially crossed across representation type and temporal handling, evaluated on a real-world production dataset. The comparison reveals how these strategies differ across user behavior types, across both accuracy and beyond-accuracy dimensions of recommendation quality, and across the temporal-window setting that governs the disentanglement.