超越借用的历史:用于交互式角色扮演评估的人物对齐用户模拟
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
浏览论文内容
中文总结 AI 辅助
本研究针对现有RPA评估基准的局限,提出基于用户模拟器的PALATE基准,通过个性化评分标准和多轮自由对话评估,生成针对特定用户-RPA对的可解释评估。
中文摘要 AI 辅助
角色扮演智能体(RPA)已成为大语言模型最重要的消费级应用之一。用户与RPA进行多轮对话以获得情感慰藉等体验,因此可靠的评估对衡量其能力、比较不同系统及指导进一步改进至关重要。然而,现有基准通常要求RPA基于固定对话历史生成后续内容,再用与用户脱节的固定评分标准评估该后续内容。我们通过实证研究指出并证明了这种设计的两个局限:第一,RPA的输出受前置对话历史影响,无法对其在真实多轮场景下的角色扮演能力进行科学严谨的评估;第二,不同个体的用户体验差异显著,传统固定评分标准未必与用户满意度一致。为此,我们提出PALATE(Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation,即采用定制化评估的人物对齐大语言模型模拟用户评估),这是一个基于用户模拟器的可扩展RPA基准。PALATE配套包含30个人物档案库,其主要评估流程为:训练5个针对每个用户的模拟器,让它们在预先固定的人物档案面板上与候选RPA进行自由形式的多轮对话。除通用质量评分标准外,我们还构建了个性化评分标准以衡量用户满意度;在保留的标注数据上,个性化评分标准与人类判断的一致性高于通用评分标准。在对16个候选模型的主要评估中,PALATE在由每个候选模型共同构建的多轮对话轨迹上,分别表征了通用轮次质量、长期会话能力以及每个用户的体验,从而生成针对特定用户-RPA对的可解释评估,而非将系统压缩为单一的、与用户无关的排名。
英文摘要
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- MetaStone Technology(MetaStone科技公司)
机构由 AI 辅助整理,请以论文原文为准。