发表机构
University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有多模态序列推荐方法忽略协同信号或计算开销大的问题,提出MGRASRec框架,通过检索用户-物品交互图的协同过滤路径注入信号,在三个公开数据集上实现最优推荐性能。
AI 中文摘要
多模态大语言模型(Multimodal Large Language Models, MLLMs)凭借对复杂多模态数据的推理能力,在序列推荐领域展现出强大潜力。然而现有方法要么仅依赖目标用户自身的交互历史,忽略了相邻用户的协同信号;要么因对长交互历史重复进行MLLM推理而产生大量计算开销。为应对这些挑战,本文提出MGRASRec,一种用于序列推荐的多模态图检索增强框架。MGRASRec通过从用户-物品交互图中检索结构化路径,将候选物品条件下的协同过滤信号直接注入MLLM的提示词中;该交互图通过多模态相似度进行扩展,以覆盖精确共同交互重叠之外的范围。这种检索还能以无额外成本的方式呈现与候选物品最相关的历史物品,无需循环总结,且每个候选物品的推理仅需一次前向传播。所有组件被统一为增强提示词,用于MLLM的参数高效微调。在三个公开数据集上的广泛评估验证了MGRASRec的有效性,其在所有指标上均取得最佳性能,尤其在排序质量方面提升显著。
英文摘要
Multimodal Large Language Models (MLLMs) have demonstrated strong potential for sequential recommendation through their ability to reason over complex multimodal data. However, existing approaches either rely solely on the target user's own interaction history, neglecting collaborative signals from neighboring users, or incur substantial computational overhead through repeated MLLM inference over long interaction histories. To address these challenges, we propose MGRASRec, a multimodal graph retrieval-augmented framework for sequential recommendation. MGRASRec injects collaborative filtering signals conditioned on the candidate item directly into the MLLM prompt by retrieving structured paths from a user-item interaction graph, extended via multimodal similarity to increase coverage beyond exact co-interaction overlap. This retrieval also surfaces the history items most relevant to the candidate at no additional cost, removing the need for recurrent summarization and keeping inference to a single forward pass per candidate. All components are unified into an augmented prompt for parameter-efficient fine-tuning of an MLLM. Extensive evaluations across three publicly available datasets validate the effectiveness of MGRASRec, achieving the best performance on all metrics with particularly strong gains in ranking quality.
CommentsAccepted to AJCAI 2026