arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29410cs.IR

智能体作为多模态推荐中的知识整合器与利用器

Agents as Knowledge Integrator and Utilizer in Multimodal Recommendation

Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Puzhen Wu, Zewei Liu, Zheng Lin, Jianheng Tang, Jing Yang, Wei Wang, Xiping Hu, Edith Ngai

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多模态推荐中信号与目标不匹配的问题,提出AgentMMRec框架,通过整合与利用两个智能体角色优化图结构与表示,在亚马逊数据集上提升了推荐指标,且适用于稀疏与冷启动场景。

中文摘要 AI 辅助

在线平台日益依赖多模态推荐系统对商品、媒体及其他网络内容进行排序。现有方法通常将视觉与文本特征注入物品表示,或基于模态级相似度构建同构图,但生成的信号仍可能与推荐目标不匹配。我们从知识整合视角研究这一语义差距:多模态内容应在用于构建推荐图或调整排序前,与用户行为共同解读。我们提出AgentMMRec,一种基于智能体的多模态推荐框架,包含两个协同角色。整合智能体(Integrator Agent)从训练交互与物品内容中推断感知用户偏好与物品属性,存储至可复用的知识记忆中。利用智能体(Utilizer Agent)在冻结的评估时记忆下,利用该记忆优化模态特定的物品-物品图、构建感知行为的同构图并重排候选列表。此设计区别于直接的LLM特征增强与纯LLM重排,因生成的知识先转换为图结构与模型表示后再用于推荐。在三个亚马逊多模态推荐数据集上的实验显示,AgentMMRec较近期多模态基线持续提升召回率(Recall)与归一化折损累计增益(NDCG),在稀疏性与物品冷启动场景下仍保持有效,且能将其构建的知识迁移至现有骨干模型。

英文摘要

Online platforms increasingly rely on multimodal recommender systems to rank products, media, and other Web content. Existing methods usually inject visual and textual features into item representations or build homogeneous graphs from modality-level similarity, but the resulting signals can remain misaligned with the recommendation objective. We study this semantic gap from a knowledge-integration perspective: multimodal content should be interpreted together with user behavior before it is used to construct recommendation graphs or adjust rankings. We propose AgentMMRec, an agent-based multimodal recommendation framework with two coordinated roles. The Integrator Agent infers behavior- and multimodal-aware user preferences and item properties from training interactions and item content, then stores them in a reusable knowledge memory. The Utilizer Agent consumes this memory to refine modality-specific item-item graphs, construct behavior-aware homogeneous graphs, and rerank candidate lists under a frozen evaluation-time memory. This design differs from direct LLM feature augmentation and pure LLM reranking because the generated knowledge is first converted into graph structure and model representations before recommendation. Experiments on three Amazon multimodal recommendation datasets show that AgentMMRec consistently improves Recall and NDCG over recent multimodal baselines, remains effective under sparsity and item cold-start settings, and can transfer its constructed knowledge to existing backbones.

发表机构

  • The University of Hong Kong(香港大学)
  • Beijing Institute of Technology(北京理工大学)
  • Peking University(北京大学)
  • Universiti Malaya(马来亚大学)
  • Macao Polytechnic University(澳门理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑