发表机构
Coller School of Management, Tel Aviv University(特拉维夫大学科勒尔管理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出从图像中提取情境信号(物理、社会、模态类别)并融合进推荐系统,通过ICE-Fuse与图对比学习验证,图像情境与评论信号互补,可增强推荐性能。
AI 中文摘要
情境信息捕捉用户-物品交互的周围环境,是推荐系统的核心。先前的工作从位置、时间或评论中提取情境,但未涉及图像;多模态推荐系统主要使用图像来丰富物品或用户表示,而非识别情境性上下文。我们提出一种从图像导出的新情境表示,涵盖通过视觉-语言模型学习的物理、社会和模态类别。我们引入ICE-Fuse,一个评估该表示的流水线,它融合这些类别并将其整合到情境感知推荐系统中,使用TripAdvisor数据和基于评论的图对比学习作为推荐算法。图像情境单独使用时并未超越既有信号,但结合使用时提升了它们,表明信息具有互补性。语义分析显示,图像和评论导出的情境捕捉交互的不同方面,将图像定位为互补情境。
英文摘要
Contextual information, capturing the circumstances of a user-item interaction, is central to recommender systems. Prior work draws context from location, time, or reviews, but not images; multimodal recommender systems mainly use images to enrich item or user representations, not identify situational context. We propose a new representation of context derived from images, spanning physical, social, and modal categories learned via a vision-language model. We introduce ICE-Fuse, a pipeline for evaluating this representation that fuses these categories and integrates them into a context-aware recommender system, using TripAdvisor data and Review-aware Graph Contrastive Learning as the recommendation algorithm. Image context does not outperform established signals standalone, but improves them combined, indicating complementary information. Semantic analysis shows image- and review-derived context capture distinct aspects of the interaction, positioning images as complementary context.
CommentsAccepted at the CARS workshop, RecSys 2026. 8 pages, 2 figures