面向多模态推荐的隐式对齐推理
Latent-Aligned Reasoning for Multimodal Recommendation
- Xiaohongshu Inc.(小红书公司)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态推荐中跨模态稀释问题,提出含互补对齐机制的两阶段隐式推理框架LARK,在多数据集上实现多种推荐架构的SOTA性能,各组件贡献经消融实验验证。
AI中文摘要:
多模态视觉语言模型(VLMs)在跨模态理解方面展现出卓越能力,但将其应用于推荐系统时存在一个核心挑战:随着表示通过多步推理传播,视觉和文本信号会逐渐衰减,我们将这一现象称为跨模态稀释。为解决该问题,我们提出LARK(Latent-Aligned Reasoning frameworK,隐式对齐推理框架),这是一个在单个VLM内具有互补对齐机制的两阶段隐式推理框架。在第一阶段,可学习的隐式标记与多步思维链(CoT)推理交织,并与冻结的视觉编码器显式对齐,作为视觉检查点在整个推理链中保留感知细节。在第二阶段,隐式表示通过桥接MLP进行投影,并使用项目间对比学习进行训练;为防止推理语义衰减,中间特征与第一阶段的CoT隐藏状态对齐,将最终嵌入锚定到模型自身的推理输出。在三个公共基准和一个工业数据集上的实验表明,LARK在多种推荐架构上实现了最先进的性能,受控 ablation( ablation即消融实验)证实了每个组件的独特贡献。
英文摘要:
Multimodal Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding, yet a fundamental challenge persists when applying them to recommendation: as representations propagate through multi-step reasoning, both visual and textual signals progressively attenuate - a phenomenon we term cross-modal dilution. To address this, we propose LARK (Latent-Aligned Reasoning frameworK), a two-stage latent reasoning framework with complementary alignment mechanisms within a single VLM. In the first stage, learnable latent tokens are interleaved with multi-step chain-of-thought (CoT) reasoning and explicitly aligned with a frozen vision encoder, serving as visual checkpoints that preserve perceptual details throughout the reasoning chain. In the second stage, the latent representations are projected via a bridge MLP and trained with item-to-item contrastive learning; to prevent the reasoning semantics from fading, intermediate features are aligned with the CoT hidden states from the first stage, anchoring the final embeddings to the model's own reasoning output. Experiments on three public benchmarks and one industrial dataset show that LARK achieves state-of-the-art performance across multiple recommendation architectures, with controlled ablations confirming the distinct contribution of each component.