arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UXBench Pro:面向多轮对话交互中个性化用户体验的基准测试

UXBench Pro: Benchmarking Personalized User Experience in Multi-Turn Dialogue Interactions

Mengze Hong, Zeyang Lei, Wenbo Shang, Xia Zeng, Xiying Zhao, Qi Zhu, Chen Jason Zhang, Di Jiang, Taiming Fu, Qiongyi Zhou, Qinghe Chang, Fubao Zhang, Chenxuan Ma, Minlong Peng, Jinfeng Huang, Zineng Zhou, Jindou Wu, Muge Qi, Sijun He, Xin Cui, Di Liang, Yuan Hua, Davey Chen

arXiv 2610.11638首次发表:更新:

发表机构

Hong Kong Polytechnic University; Tencent(香港理工大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出UXBench Pro基准测试,结合个性化用户奖励模型与用户模拟器Sim4Eval,通过URMBench、USimBench评估模型可靠性,为多轮对话的个性化用户体验评估提供新视角。

AI 中文摘要

通过自动化计算方法评估用户体验(UX)已受到越来越多关注,UXBench的实证研究为此提供了支持。然而,二元偏好预测能提供的见解有限,且依赖单一的用户无关奖励模型会忽略用户的固有异质性,不同用户的期望可能存在显著差异。本文提出UXBench Pro,包含1000个测试实例,这些实例源自12种任务场景、82个领域的真实用户交互。每个实例都配有FACTORS用户画像,该画像通过7个可解释的行为维度刻画用户,以区分不同用户群体。为提供更丰富的评估见解,本文引入双视角范式,结合用于第三人称判断的个性化用户奖励模型(URM),以及支持多轮交互、并能从四个认知状态维度提供第一人称评估的用户模拟器Sim4Eval。为评估这些基于模型的评估器的可靠性,本文进一步提出两个元基准:URMBench和USimBench,用于评估它们复现真实人类偏好和行为的忠实度。大量实验揭示了7项关键发现,凸显了用户建模和多视角评估的重要性,为以用户为中心的基准测试提供了新视角,并推动了个性化模型优化。

英文摘要

Evaluating user experience (UX) with automated computational methods has gained increasing attention, supported by empirical evidence from UXBench. However, binary preference prediction provides limited insight, while relying on a single user-agnostic reward model overlooks the inherent heterogeneity of users, whose expectations can differ substantially. In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains. Each instance is paired with a FACTORS user profile that characterizes the user through seven interpretable behavioral facets, differentiating user groups. To provide richer evaluation insights, we introduce a dual-perspective paradigm that combines a personalized User Reward Model (URM) for third-person judgment with Sim4Eval, a user simulator that enables multi-turn interactions and provides first-person evaluation across four cognitive state dimensions. To assess the reliability of these based evaluators, we further introduce two meta-benchmarks, URMBench and USimBench, that evaluate how faithfully they reproduce real human preferences and behaviors. Extensive experiments reveal seven key findings that highlight the importance of user modeling and multi-perspective evaluation, offering a fresh perspective on user-centric benchmarking and motivating personalized model optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑