arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ComboShoppingBench:评估大型语言模型智能体在带优惠券的预算约束组合购物篮任务中的表现

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

Adrian Li, Kelong Mao, Yudong Guo, Heming Xia, Xinwei Yang, Lirui Luo, Jace Wong, Pu Yao, Sulong Xu, Simiu Gu

arXiv 2608.09282首次发表:更新:

AI 中文总结

本研究推出ComboShoppingBench组合购物基准,通过探索智能体、LLM评判者及确定性验证,发现各类LLM智能体在该基准上表现不佳,凸显其在约束感知组合购物上仍有巨大改进空间。

AI 中文摘要

现实购物常需要构建互补商品的购物篮,而非仅检索单一产品,这类组合购物任务出现在设备配置、餐食准备、活动规划、团体外卖订购等场景,需联合推理商品兼容性、库存、店铺要求、配送费、优惠券及预算。评估极具挑战性,因为多个购物篮可能满足同一请求, exact-match指标不适用,而仅语义评估无法检测不可行订单、无效优惠券组合或错误支付。我们推出ComboShoppingBench,这是一个用于在模拟商业和外卖环境中构建开放式但可验证购物篮的智能购物基准。任务合成阶段,探索智能体构建可行且语义连贯的可购买商品篮,该“见证者”引导生成优惠券、预算约束、用户查询及对齐的评估标准;评估阶段,LLM评判者评估语义满意度、响应质量和声明忠实度,同时确定性验证检查商品ID有效性、预算合规性和优惠券最优性。对各类LLM智能体的实验表明,即使是强大的智能体在ComboShoppingBench上也表现不佳,凸显了可靠、感知约束的组合购物方面仍有巨大改进空间。

英文摘要

Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, requiring joint reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets. Evaluation is challenging because multiple baskets may satisfy the same request, making exact-match metrics unsuitable, whereas semantic evaluation alone cannot detect infeasible orders, invalid coupon combinations, or incorrect payments. We introduce ComboShoppingBench, an agentic shopping benchmark for open-ended yet verifiable basket construction in a simulated commerce and takeout environment. During task synthesis, an exploration agent constructs a feasible and semantically coherent basket of purchasable products; this witness guides the generation of coupons, budget constraints, user queries, and aligned evaluation rubrics. During evaluation, LLM judges assess semantic satisfaction, response quality, and claim faithfulness, while deterministic validation checks product-ID validity, budget compliance, and coupon optimality. Experiments with diverse LLM agents demonstrate that even strong agents struggle on ComboShoppingBench, highlighting substantial room for improvement in reliable, constraint-aware combo shopping.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑