arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PersonaEval:用于评估交互式应用的基于角色的用户模拟

PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications

Yifan Simon Liu, Qianfeng Wen, Yilan Fan, Shirley Huang, Ruoqi Gao, Jianheng Hou, Muhammad Ahmed Mohsin, Zonglin Di, Brihi Joshi, Xincheng Tan, Yucheng Lu, Xiaoyi Liu, Heming Liu, Hanwen Xing, Guanghui Min, Zhengyang Shan, My Chiffon Nguyen, Ishan Gupta, Yunze Xiao, Hannah Collison, Jintao Huang, Jiatong Li, Sankalp Jajee, Yunhan Zhao, Bing Hu, Sky Ng, Xupeng Chen, Binghang Lu, Weihang Xiao, Aravind Mohan, Bolun Sun, Yunshu Wu, Yuanda Xu, Yun Shen, Runyu Zhang, Zheyuan Deng, Zhiwei Zhang, Qianyu Zhu, Dianzhuo Wang, Yijun Wang, Yixuan He, Yuexing Hao, Xiaomin Li

arXiv 2608.15838首次发表:更新:

AI 中文总结

针对真实用户研究成本高、难扩展的问题,提出PersonaEval框架,可在三类交互式应用上实现可重复、可并行的用户模拟评估。

AI 中文摘要

真实用户研究对于理解人们与被测或已部署系统的交互方式十分重要,但实际开展时往往成本高、耗时长且难以扩展。为应对这些挑战,我们引入PersonaEval,这是一个基于角色的用户模拟框架,可在各类交互式场景中近似模拟真实用户行为。PersonaEval将来自现有角色数据集的模拟用户与特定任务的应用界面相连接,收集交互轨迹与结果。该框架提供即插即用的评估工作流,可轻松更换待评估应用。本次演示中,我们展示了PersonaEval在调查、聊天机器人、网页应用这三类交互式应用上的应用。这些示例表明,PersonaEval可支持在不同交互场景下开展可重复、可并行、可扩展的评估,同时生成面向用户的反馈及特定任务的行为。

英文摘要

Real user studies are important for understanding how people interact with systems under test or already deployed. In practice, however, they are often costly, time-consuming, and difficult to scale. To address these challenges, we introduce PersonaEval, a persona-based user simulation framework that approximates real-user behavior across diverse interactive settings. PersonaEval connects simulated users drawn from existing persona datasets to task-specific application interfaces and collects the interaction trajectories and outcomes. PersonaEval provides a plug-and-play evaluation workflow in which the application being evaluated can be easily changed. In this demo, we present PersonaEval on three forms of interactive applications: surveys, chatbots, and web applications. Together, these examples show that PersonaEval can support repeatable, parallelizable, and scalable evaluation across different interaction settings, while producing user-oriented feedback and task-specific behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑