发表机构
Stanford University; Google DeepMind; University of Southern California; Amazon; Massachusetts Institute of Technology; MIT Sloan School of Management(斯坦福大学; 谷歌DeepMind; 南加州大学; 亚马逊公司; 麻省理工学院; 麻省理工学院斯隆管理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对固定目录场景下会话式电商推荐的可靠性问题,提出混合多智能体框架MACS,其在单轮、多轮基准测试中均展现出更强的约束合规性与推荐性能。
AI 中文摘要
电商领域的会话式推荐越来越多地由大语言模型(LLM)驱动,但许多实际部署场景有更严格的要求:推荐必须仅来自商家的固定商品目录,不能使用网络搜索或未经支持的商品声明。在此场景下,核心挑战是硬约束下的可靠性:系统必须满足用户需求、严格基于现有库存、并在多轮会话中保留用户偏好。我们提出MACS(Multi-Agent Commerce System,多智能体电商系统),一种面向固定目录场景下可靠会话式推荐的混合多智能体框架。MACS将大语言模型用于语言相关任务,如解析用户请求、挖掘偏好、生成回复;而对正确性至关重要的操作,包括商品检索、硬约束过滤、品牌排除、渐进式放松,则由商家智能体确定性执行。会话持久偏好层跟踪多轮会话中的约束,支持预算覆盖和排除反转的一致处理。在包含140个查询的单轮基准测试中,MACS达到最高通过率(87.1%)和完美品牌合规性(1.000);在包含10个场景的多轮基准测试中,MACS达到最强的宏Pass@5(72%,对比GPT+Catalog的56%、Gemini+Catalog的52%),且零约束漂移,其优势在排除反转(100%对比20%、0%)和约束累积(100%对比60%、40%)时最为显著,各系统的平均人工判定回复质量相近(0.751对比0.736)。这些结果表明,结合确定性约束执行与会话持久偏好跟踪的混合架构,在固定目录商家场景中,比仅基于目录的纯提示基线系统提供更强的面向可靠性的性能。
英文摘要
Conversational recommendation for e-commerce is increasingly mediated by large language models (LLMs), yet many real-world deployments operate under a stricter requirement: recommendations must be drawn only from a merchant's fixed catalog, without web search or unsupported product claims. In this setting, the main challenge is reliability under hard constraints: the system must satisfy user requirements, remain grounded in available inventory, and preserve preferences across multiple conversational turns. We present MACS (Multi-Agent Commerce System), a hybrid multi-agent framework for reliable conversational recommendation in fixed-catalog settings. MACS uses LLMs for language-facing tasks such as interpreting user requests, eliciting preferences, and generating responses, while correctness-critical operations, including product retrieval, hard-constraint filtering, brand exclusion, and progressive relaxation, are executed deterministically by the merchant agent. A session-persistent preference layer tracks constraints across turns, enabling consistent handling of budget overwrites and exclusion reversals. On a 140-query single-turn benchmark, MACS achieves the highest pass rate (87.1%) and perfect brand compliance (1.000). On a 10-scenario multi-turn benchmark, MACS achieves the strongest macro Pass@5 (72% vs. 56% GPT+Catalog / 52% Gemini+Catalog) with zero constraint drift. The advantage is sharpest on exclusion reversal (100% vs. 20% / 0%) and constraint accumulation (100% vs. 60% / 40%). Mean judged response quality is similar across systems (0.751 vs. 0.736). These results suggest that hybrid architectures combining deterministic constraint enforcement with session-persistent preference tracking provide stronger reliability-oriented performance than catalog-bound prompt-only baselines in the fixed-catalog merchant setting.
Comments9 pages, 2 figures, 8 tables. Will be presenting at Stanford Trust&Safety Conference, already presented at Stanford Market AI Conference