PACEShop:评估个性化、可操作、可组合且基于证据的购物助手
PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants
浏览论文内容
中文总结 AI 辅助
针对现有购物助手评估基准未覆盖结构化回复联合评估目标的问题,提出 PACE 评估框架,构建含 22625 条记录的 PACEShop 基准数据集和 PACEJudge 协议,验证任务匹配输出契约对评估的重要性。
中文摘要 AI 辅助
购物助手正从排序的产品列表转向结构化决策支持,系统必须将购物者情境、产品证据和下一步指导综合为连贯的推荐体验。这改变了评估单位:流畅的回复仍可能因忽略购物者情境、组件间自相矛盾或缺陷过于模糊而无法定位而失败。现有的个性化、基于证据和 LLM 作为评判者的基准仅覆盖该问题的部分内容,但未定义结构化购物助手回复的联合评估目标。我们将这一缺失的评估目标定义为 PACE:即个性化(Personalized)、可操作(Actionable)、可组合(Compositional)且基于证据(Evidence-grounded)的评估。我们通过两个实例实现 PACE:一是 PACEShop,这是一个基准数据集,通过 22625 条受控记录使该目标可测量,这些记录包含结构化角色、可审计的证据池、GOOD/BAD 标签以及缺陷类别和位置的黄金标注;二是 PACEJudge,这是一种无需训练的评判协议,通过结构化输出契约使该目标可报告。我们的实验表明,通用评判者能够识别广泛的质量问题,但无法恢复 PACE 所需的诊断字段;PACEShop 使这些失败可验证,而 PACEJudge 无需重新训练即可提升角色来源、跨组件、基于证据以及类别和位置的闭合度,这表明现实的购物助手评估需要与任务匹配的输出契约,而非仅依赖更强的骨干模型或标量提示。
英文摘要
Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This changes the unit of evaluation: a fluent response can still fail by ignoring shopper context, contradicting itself across components, or leaving defects too vague to localize. Existing personalization, grounding, and LLM-as-a-judge benchmarks cover pieces of this problem, but they do not define a joint evaluation target for structured shopping-assistant responses. We formulate this missing evaluation target as PACE: Personalized, Actionable, Compositional, and Evidence-grounded evaluation. We instantiate PACE with two artifacts: PACEShop, a benchmark dataset that makes the target measurable through 22,625 controlled records with structured personas, auditable evidence pools, GOOD/BAD labels, and gold defect family and location annotations; and PACEJudge, a training-free judging protocol that makes the target reportable through a structured output contract. Our experiments show that generic judges can recognize broad quality but fail to recover the diagnostic fields required for PACE; PACEShop makes these failures verifiable, and PACEJudge improves persona-source, cross-component, grounding, and family/location closure without retraining, showing that realistic shopping-assistant evaluation requires a task-matched output contract rather than only a stronger backbone or scalar prompt.
发表机构
- Amazon(亚马逊)
- Northeastern University(东北大学)
- Expedia(亿客行)
机构由 AI 辅助整理,请以论文原文为准。