何时依赖视觉信息:序列推荐中的门控多模态融合
Deciding When to Rely on Visual Information: Gated Multimodal Fusion in Sequential Recommendation
AI总结:
该研究提出VisGate框架,基于物品和用户上下文自适应融合多模态信号,揭示视觉效用随物品、交互稀疏性变化的规律,实现更精准的序列推荐并可解释视觉信息的作用。
AI中文摘要:
多模态序列推荐系统通常将视觉信号与协同信号统一融合,无论物品或用户上下文如何,都将视觉特征视为通用信息。我们认为,视觉效用(定义为视觉信号对推荐质量的贡献)是一种潜在的上下文变量,取决于物品和用户的交互历史,而非固定的物品属性。为了建模这种变异性,我们引入VisGate框架,该框架基于物品嵌入和用户当前序列上下文做出自适应的物品级融合决策。视觉表示通过序列共现模式上的对比目标学习,保留与协同嵌入的互补性,而非将其对齐到共享空间。除了实现有竞争力的推荐性能外,VisGate学习到的门还作为一种测量工具,用于理解视觉信息何时以及为何有益。我们的分析表明,视觉效用在不同物品间存在差异,在交互稀疏性下(当协同信号较弱时)会增加,并以语义有意义的方式与视觉独特性相关联。这些发现共同强调了细粒度融合和模态互补性的重要性,同时证明物品级视觉效用可通过学习到的门控行为进行估计和解释。
英文摘要:
Multimodal sequential recommender systems commonly fuse visual and collaborative signals uniformly, treating visual features as generically informative regardless of item or user context. We argue that visual utility, defined as the contribution of visual signals to recommendation quality, is a latent contextual variable that depends on both the item and the user's interaction history rather than a fixed item property. To model this variability, we introduce VisGate, a framework that makes adaptive item-level fusion decisions conditioned on item embeddings and the user's current sequence context. Visual representations are learned through a contrastive objective over sequential co-occurrence patterns, preserving complementarity with collaborative embeddings rather than aligning them into a shared space. Beyond achieving competitive recommendation performance, VisGate's learned gate serves as a measurement tool for understanding when and why visual information is beneficial. Our analyses show that visual utility varies across items, increases under interaction sparsity when collaborative signals are weak, and correlates with visual distinctiveness in semantically meaningful ways. Together, these findings highlight the importance of both fine-grained fusion and modality complementarity, while demonstrating that item-level visual utility can be estimated and interpreted through learned gating behaviour.