arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习多模态内容的稀疏表示以增强冷启动项目推荐

Learning Sparse Representations of Multimodal Content for Enhanced Cold Item Recommendation

Gregor Meehan, Johan Pauwels

arXiv 2607.17184首次发表:更新:

AI 中文总结

针对数字平台项目目录规模大及冷启动问题,提出利用稀疏嵌入优势,通过改进冷启动训练方法和设计预稀疏化激活技术,降低存储成本并提升冷启动推荐准确性,还展示了稀疏内容嵌入的可解释性与稳健性。

AI 中文摘要

现代数字平台中项目目录的规模和快速增长给推荐系统从业者带来了重大挑战。大多数推荐系统使用嵌入相似度来预测用户-项目偏好,但在行业规模的目录中,嵌入存储和低延迟检索具有挑战性。此外,新添加的项目没有相应的嵌入,无法有效推荐。以往的工作通常通过从辅助内容(如图像或描述性文本)生成冷启动项目表示来解决这个问题。本文认为,在基于内容的冷启动范式中,稀疏嵌入比标准密集向量具有显著优势。我们描述了如何将现有的冷启动训练方法适用于稀疏表示学习,并基于线性注意力的见解设计了一种预稀疏化激活技术,该技术在学习到的项目-项目相似度中诱导锐化和去噪效果。我们表明,所得的稀疏嵌入在冷启动推荐准确性方面比密集嵌入有显著提高,同时存储成本大大降低,特别是对于具有多个兴趣的用户。通过对四个多模态推荐系统数据集的综合实验我们还展示了稀疏内容嵌入的可解释性及其在大小和准确性之间权衡的稳健性。

英文摘要

The scale and rapid growth of item catalogs in modern digital platforms present significant challenges to recommender system (RS) practitioners. Most RSs use embedding similarity to predict user-item preferences, but embedding storage and low-latency retrieval are challenging in industry-scale catalogs. Furthermore, newly added items do not have corresponding embeddings and cannot be recommended effectively; previous works often tackle this item cold-start problem by generating cold item representations from auxiliary content, such as images or descriptive text, so that user preferences can be predicted without historical interactions. In this paper, we argue that sparse embeddings have notable advantages over standard dense vectors in this content-based cold-start paradigm. We describe how existing cold-start training regimes can be adapted for sparse representation learning, and build on insights from linear attention to design a pre-sparsification activation technique that induces sharpness and denoising effects in learned item-item similarities. We show that the resulting sparse embeddings achieve significant improvements in cold-start recommendation accuracy over dense embeddings at considerably lower storage costs, especially for users with multiple interests. Through comprehensive experiments on four multimodal RS datasets, we also demonstrate the interpretability of sparse content embeddings and their robustness in the trade-off between size and accuracy.

CommentsAccepted at RecSys 2026

DOI:10.1145/3773078.3831760

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑