OpticalRec:用于多模态推荐系统的统一光学视觉-语言表示
OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation
浏览论文内容
中文总结 AI 辅助
OpticalRec提出首个视觉空间统一编码范式,通过将文本元数据渲染为视觉字形,利用视觉-语言模型的双重注意力机制,实现多模态协同过滤中的原生图像-文本交互,提升推荐性能并支持即插即用。
中文摘要 AI 辅助
近年来,视觉-语言建模的进展显著提升了多模态编码、检索和推理的能力。然而,在多模态推荐系统中,编码丰富的物品视觉-语言语义交互仍是一个长期存在的瓶颈,这阻碍了准确的物品表示学习和用户-物品匹配。主流方法主要采用视觉和语言模态的独立编码,随后进行刚性后期融合(如拼接),这本质上忽略了原生的视觉-语言交互,并引入了跨模态语义失真。为解决这一挑战,我们提出了OpticalRec,这是首个用于多模态协同过滤(一种基础推荐设置)的视觉空间统一编码范式。OpticalRec不是采用孤立的模态特定编码,而是将物品的文本元数据渲染为视觉字形,使视觉编码器(感知编码层)内能够进行原生的图像-文本交互。由此产生的表示进一步由语言解码器(语义编码层)处理,使OpticalRec能够利用现代视觉-语言模型的双重注意力机制,而之前的编码方法忽略了这一机制。OpticalRec的有效性在理论上得到了互信息分析的支持,并在强基线和基准上通过卓越的性能得到了实证证明。作为一个即插即用模块,OpticalRec(1)引入了极低的成本,(2)对渲染文本的字体、颜色和布局等具有鲁棒性,(3)能无缝集成到现有的多模态协同过滤模型中。
英文摘要
Recent advances in vision-language modeling have substantially improved multimodal encoding, retrieval and reasoning. Yet for multimodal recommendation, encoding rich item vision-language semantic interactions remains a long-standing bottleneck, which hampers accurate item representation learning and user-item matching. Mainstream approaches primarily adopt independent encoding of vision and language modality followed by rigid late fusion such as concatenation, inherently omitting native vision-language interactions and introducing cross-modal semantic distortion. To address this challenge, we propose OpticalRec, the first visual-space unified encoding paradigm for multimodal collaborative filtering, a fundamental recommendation setting. Instead of isolated modality-specific encoding, OpticalRec renders item textual metadata as visual glyphs, enabling native image-text interaction within the visual encoder - the perceptual encoding level. The resulting representations are further processed by the language decoder - the semantic encoding level, allowing OpticalRec to exploit the dual-attention mechanism of modern vision-language models that previous encoding methods omitted. OpticalRec's efficacy is theoretically supported by mutual information analysis and empirically demonstrated through superior performance across strong baselines and benchmarks. As a plug-and-play module, OpticalRec (1) introduces minimal cost, (2) is robust against rendered text font, color and layout, etc., and (3) integrates seamlessly into existing multimodal collaborative filtering models.
发表机构
- University of California, San Diego(加州大学圣地亚哥分校)
- Alibaba Group(阿里巴巴集团)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Texas A&M University(德克萨斯农工大学)
- The Hong Kong Polytechnic University(香港理工大学)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。