AI 中文总结
该研究推出首个多模态多轮购物智能体真实日志基准MMShopBench,构建离线购物沙箱,经微调后开源模型性能大幅接近领先专有模型,验证了基准及配套训练数据的有效性。
AI 中文摘要
在线购物者越来越多地使用AI购物助手,借助图像和多轮对话来表达和细化仅用文字难以清晰表述的产品需求。然而,现有基准大多依赖纯文本或合成请求,未能充分体现通过图像与语言共同表达的复杂现实购物需求。我们推出MMShopBench,这是首个面向多模态多轮购物智能体的真实日志基准。该基准由精心清洗并人工标注的购物日志构建,提供每个请求的购买意图及强制性产品需求的真实标注。智能体需从用户图像与多轮对话中联合推断这些需求,通过图像和文本搜索检索候选产品,并利用产品图像及结构化属性验证每个候选是否满足所有需求。我们采用基于证据的多模态协议评估代表性开源和专有模型,并构建配套训练集以微调开源模型。为确保实验可复现,我们搭建了离线购物沙箱,其中微调操作大幅缩小了我们的开源模型与领先专有模型之间的性能差距,证明了我们训练数据的有效性。
英文摘要
Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic requests, underrepresenting complex real-world shopping requirements jointly expressed through images and language. We introduce MMShopBench, the first real-log benchmark for multimodal, multi-turn shopping agents. Built from carefully cleaned and manually annotated shopping logs, MMShopBench provides ground-truth annotations of each request's purchase intent and mandatory product requirements. Agents must infer these requirements jointly from user images and multi-turn dialogue, retrieve candidate products through image and text search, and verify that each candidate satisfies all requirements using its product images and structured attributes. We evaluate representative open-source and proprietary models using an evidence-grounded multimodal protocol and construct a companion training set for fine-tuning an open-source model. To ensure reproducible experimentation, we build an offline shopping sandbox, where fine-tuning substantially narrows the performance gap between our open-source model and leading proprietary models, demonstrating the effectiveness of our training data.
Comments16 pages, 6 figures, including appendix