arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26180cs.CLcs.AI

PACEShop:评估个性化、可操作、可组合且基于证据的购物助手

PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants

Weimin Lyu, Chen Luo, Guangrui Li, Yaochen Xie, Dhineshkumar Ramasubbu, Arief Koesdwiady, Wanqiu Long, Hansu Gu, Yutong Chen, Zheshen Wang, Dakuo Wang, Yi Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有购物助手评估基准未覆盖结构化回复联合评估目标的问题,提出 PACE 评估框架,构建含 22625 条记录的 PACEShop 基准数据集和 PACEJudge 协议,验证任务匹配输出契约对评估的重要性。

中文摘要 AI 辅助

购物助手正从排序的产品列表转向结构化决策支持,系统必须将购物者情境、产品证据和下一步指导综合为连贯的推荐体验。这改变了评估单位:流畅的回复仍可能因忽略购物者情境、组件间自相矛盾或缺陷过于模糊而无法定位而失败。现有的个性化、基于证据和 LLM 作为评判者的基准仅覆盖该问题的部分内容,但未定义结构化购物助手回复的联合评估目标。我们将这一缺失的评估目标定义为 PACE:即个性化(Personalized)、可操作(Actionable)、可组合(Compositional)且基于证据(Evidence-grounded)的评估。我们通过两个实例实现 PACE:一是 PACEShop,这是一个基准数据集,通过 22625 条受控记录使该目标可测量,这些记录包含结构化角色、可审计的证据池、GOOD/BAD 标签以及缺陷类别和位置的黄金标注;二是 PACEJudge,这是一种无需训练的评判协议,通过结构化输出契约使该目标可报告。我们的实验表明,通用评判者能够识别广泛的质量问题,但无法恢复 PACE 所需的诊断字段;PACEShop 使这些失败可验证,而 PACEJudge 无需重新训练即可提升角色来源、跨组件、基于证据以及类别和位置的闭合度,这表明现实的购物助手评估需要与任务匹配的输出契约,而非仅依赖更强的骨干模型或标量提示。

英文摘要

Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This changes the unit of evaluation: a fluent response can still fail by ignoring shopper context, contradicting itself across components, or leaving defects too vague to localize. Existing personalization, grounding, and LLM-as-a-judge benchmarks cover pieces of this problem, but they do not define a joint evaluation target for structured shopping-assistant responses. We formulate this missing evaluation target as PACE: Personalized, Actionable, Compositional, and Evidence-grounded evaluation. We instantiate PACE with two artifacts: PACEShop, a benchmark dataset that makes the target measurable through 22,625 controlled records with structured personas, auditable evidence pools, GOOD/BAD labels, and gold defect family and location annotations; and PACEJudge, a training-free judging protocol that makes the target reportable through a structured output contract. Our experiments show that generic judges can recognize broad quality but fail to recover the diagnostic fields required for PACE; PACEShop makes these failures verifiable, and PACEJudge improves persona-source, cross-component, grounding, and family/location closure without retraining, showing that realistic shopping-assistant evaluation requires a task-matched output contract rather than only a stronger backbone or scalar prompt.

发表机构

  • Amazon(亚马逊)
  • Northeastern University(东北大学)
  • Expedia(亿客行)

机构由 AI 辅助整理,请以论文原文为准。

↑