RealWorldShop:真实世界电子商务中对话式购物智能体的基准测试与改进
RealWorldShop: Benchmarking and Improving Conversational Shopping Agents in Real-World E-commerce
- College of Computer Science, Sichuan University(四川大学计算机学院)
- Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对真实电商购物中现有基准忽视会话级决策过程的问题,提出REALWORLDSHOP基准及REALSHOP_AGENT会话控制框架,通过显式状态管理等方法显著提升性能。
AI中文摘要:
大语言模型正在将电子商务从静态推荐器重塑为交互式购物助手,然而真实世界的购物需要会话级决策支持:用户会揭示并修正约束条件,协调多个目标,并期望在整个对话中获得基于产品的推荐。现有基准大多面向结果或执行导向,使得这一不断演进的决策过程评估不足。我们引入了REALWORLDSHOP,一个基于328万条落地产品的基准,包含结构化的购物会话、基于用户画像且动作受控的用户模拟器,以及角色扮演评估。我们的分析表明,当前系统能产生局部合理的响应,但在状态跟踪、约束更新和落地收敛方面存在困难,尤其是在意图模糊、捆绑购买和多意图场景下。我们进一步提出了REALSHOP_AGENT,一个可执行的会话控制框架,具备显式状态管理、购物流程控制、基于目录的检索和运行时防护。实验表明,REALSHOP_AGENT在REALWORLDSHOP上持续优于强基线模型。
英文摘要:
Large language models are reshaping ecommerce from static recommenders into interactive shopping assistants, yet real-world shopping requires session-level decision support: users reveal and revise constraints, coordinate multiple goals, and expect product-grounded recommendations over a full conversation. Existing benchmarks are mostly outcome-oriented or execution-oriented, leaving this evolving decision process under-evaluated. We introduce REALWORLDSHOP, a benchmark built on 3.28M grounded products, structured shopping episodes, a profile-grounded and actioncontrolled user simulator, and role-play evaluation. Our analysis shows that current systems produce locally plausible responses but struggle with state tracking, constraint updating, and grounded convergence, especially under ambiguous intent, bundle, and multi-intent scenarios. We further propose REALSHOP_AGENT, an executable session-control framework with explicit state management, shopping-flow control, catalog-grounded retrieval, and runtime guards. Experiments show that REALSHOP_AGENT consistently outperforms strong baselines on REALWORLDSHOP.