发表机构
Hong Kong Polytechnic University; Tencent(香港理工大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出UXBench Pro基准测试,结合个性化用户奖励模型与用户模拟器Sim4Eval,通过URMBench、USimBench评估模型可靠性,为多轮对话的个性化用户体验评估提供新视角。
AI 中文摘要
通过自动化计算方法评估用户体验(UX)已受到越来越多关注,UXBench的实证研究为此提供了支持。然而,二元偏好预测能提供的见解有限,且依赖单一的用户无关奖励模型会忽略用户的固有异质性,不同用户的期望可能存在显著差异。本文提出UXBench Pro,包含1000个测试实例,这些实例源自12种任务场景、82个领域的真实用户交互。每个实例都配有FACTORS用户画像,该画像通过7个可解释的行为维度刻画用户,以区分不同用户群体。为提供更丰富的评估见解,本文引入双视角范式,结合用于第三人称判断的个性化用户奖励模型(URM),以及支持多轮交互、并能从四个认知状态维度提供第一人称评估的用户模拟器Sim4Eval。为评估这些基于模型的评估器的可靠性,本文进一步提出两个元基准:URMBench和USimBench,用于评估它们复现真实人类偏好和行为的忠实度。大量实验揭示了7项关键发现,凸显了用户建模和多视角评估的重要性,为以用户为中心的基准测试提供了新视角,并推动了个性化模型优化。
英文摘要
Evaluating user experience (UX) with automated computational methods has gained increasing attention, supported by empirical evidence from UXBench. However, binary preference prediction provides limited insight, while relying on a single user-agnostic reward model overlooks the inherent heterogeneity of users, whose expectations can differ substantially. In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains. Each instance is paired with a FACTORS user profile that characterizes the user through seven interpretable behavioral facets, differentiating user groups. To provide richer evaluation insights, we introduce a dual-perspective paradigm that combines a personalized User Reward Model (URM) for third-person judgment with Sim4Eval, a user simulator that enables multi-turn interactions and provides first-person evaluation across four cognitive state dimensions. To assess the reliability of these based evaluators, we further introduce two meta-benchmarks, URMBench and USimBench, that evaluate how faithfully they reproduce real human preferences and behaviors. Extensive experiments reveal seven key findings that highlight the importance of user modeling and multi-perspective evaluation, offering a fresh perspective on user-centric benchmarking and motivating personalized model optimization.