arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05432cs.IR

OpticalRec:用于多模态推荐系统的统一光学视觉-语言表示

OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation

Yueqi Wang, Zitian Guo, Yupeng Hou, Yifei Wang, Kibum Kim, Zhenrui Yue, Shuo Xing, Haodong Li, Heming Xia, Renrui Zhang, Zhengzhong Tu, Julian McAuley

首次发表
浏览论文内容

中文总结 AI 辅助

OpticalRec提出首个视觉空间统一编码范式,通过将文本元数据渲染为视觉字形,利用视觉-语言模型的双重注意力机制,实现多模态协同过滤中的原生图像-文本交互,提升推荐性能并支持即插即用。

中文摘要 AI 辅助

近年来,视觉-语言建模的进展显著提升了多模态编码、检索和推理的能力。然而,在多模态推荐系统中,编码丰富的物品视觉-语言语义交互仍是一个长期存在的瓶颈,这阻碍了准确的物品表示学习和用户-物品匹配。主流方法主要采用视觉和语言模态的独立编码,随后进行刚性后期融合(如拼接),这本质上忽略了原生的视觉-语言交互,并引入了跨模态语义失真。为解决这一挑战,我们提出了OpticalRec,这是首个用于多模态协同过滤(一种基础推荐设置)的视觉空间统一编码范式。OpticalRec不是采用孤立的模态特定编码,而是将物品的文本元数据渲染为视觉字形,使视觉编码器(感知编码层)内能够进行原生的图像-文本交互。由此产生的表示进一步由语言解码器(语义编码层)处理,使OpticalRec能够利用现代视觉-语言模型的双重注意力机制,而之前的编码方法忽略了这一机制。OpticalRec的有效性在理论上得到了互信息分析的支持,并在强基线和基准上通过卓越的性能得到了实证证明。作为一个即插即用模块,OpticalRec(1)引入了极低的成本,(2)对渲染文本的字体、颜色和布局等具有鲁棒性,(3)能无缝集成到现有的多模态协同过滤模型中。

英文摘要

Recent advances in vision-language modeling have substantially improved multimodal encoding, retrieval and reasoning. Yet for multimodal recommendation, encoding rich item vision-language semantic interactions remains a long-standing bottleneck, which hampers accurate item representation learning and user-item matching. Mainstream approaches primarily adopt independent encoding of vision and language modality followed by rigid late fusion such as concatenation, inherently omitting native vision-language interactions and introducing cross-modal semantic distortion. To address this challenge, we propose OpticalRec, the first visual-space unified encoding paradigm for multimodal collaborative filtering, a fundamental recommendation setting. Instead of isolated modality-specific encoding, OpticalRec renders item textual metadata as visual glyphs, enabling native image-text interaction within the visual encoder - the perceptual encoding level. The resulting representations are further processed by the language decoder - the semantic encoding level, allowing OpticalRec to exploit the dual-attention mechanism of modern vision-language models that previous encoding methods omitted. OpticalRec's efficacy is theoretically supported by mutual information analysis and empirically demonstrated through superior performance across strong baselines and benchmarks. As a plug-and-play module, OpticalRec (1) introduces minimal cost, (2) is robust against rendered text font, color and layout, etc., and (3) integrates seamlessly into existing multimodal collaborative filtering models.

发表机构

  • University of California, San Diego(加州大学圣地亚哥分校)
  • Alibaba Group(阿里巴巴集团)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Texas A&M University(德克萨斯农工大学)
  • The Hong Kong Polytechnic University(香港理工大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑