发表机构
ByteDance; The University of Melbourne; AXON; University of Technology Sydney(字节跳动; 墨尔本大学; AXON; 悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对在线购物个性化会话式购物的长时程偏好一致性问题,提出多智能体多模态RAG框架,经实验验证其能显著提升交互质量与用户感知个性化。
AI 中文摘要
个性化会话式购物需要在多轮交互中保持偏好一致性,用户会逐步透露约束条件。现有方法通常依赖静态用户画像,未明确控制长时程交互行为。我们提出一种多智能体、多模态的检索增强生成(RAG)框架,该框架将对话状态跟踪、推荐检索、偏好感知推理和响应生成进行分解,同时整合产品元数据、产品评论、图像衍生描述以及用户历史评论。为评估交互层面的质量,我们采用包含四个维度的轨迹级协议:全局偏好一致性、累积信息综合、交互轨迹以及语气一致性。在Amazon Reviews 2023基准上,启用检索的变体在自动轨迹指标上优于无RAG的基线(平均4.82对比3.74)。在小型真实用户研究中(n=5),完整变体获得最高平均总体评分(4.60对比基线的2.20),为角色分解加以用户为中心的检索可提升感知个性化提供了探索性证据。
英文摘要
Personalized conversational shopping requires maintaining preference consistency over multi-turn interactions, where users reveal constraints gradually. Existing approaches often rely on static profiles and do not explicitly control long-horizon interaction behavior. We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews. To evaluate interaction-level quality, we adopt a trajectory-level protocol with four dimensions: Global Preference Consistency, Cumulative Information Synthesis, Interaction Trajectory, and Tone Consistency. On an Amazon Reviews 2023 benchmark, retrieval-enabled variants outperform a no-RAG baseline on automatic trajectory metrics (average 4.82 vs. 3.74). In a small real-user study ($n{=}5$), the Full variant achieves the highest mean overall rating (4.60 vs. 2.20 for Baseline), providing exploratory evidence that role decomposition plus user-centric retrieval improves perceived personalization.\footnote{Code and dataset are available at: https://github.com/RenaGao/Multimodel_RAG_Indexing
Commentsaccepted by AACL-IJCNLP 2026