RoleRMBench & RoleRM:迈向基于角色扮演的对话系统奖励建模
RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
- Shanghai Jiao Tong University(上海交通大学)
- Fudan University(复旦大学)
- Saarland University(萨尔兰州大学)
- Tencent Youtu Lab(腾讯优图实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
RoleRMBench与RoleRM旨在通过连续隐式偏好训练提升角色扮演对话系统中奖励建模的准确性与连贯性。
AI中文摘要:
奖励建模已成为对齐大型语言模型(LLMs)与人类偏好的重要基石。然而,当扩展到主观和开放性领域,如角色扮演时,现有奖励模型表现出严重的退化,难以捕捉细腻且基于角色的人类判断。为解决这一差距,我们引入RoleRMBench,这是首个系统性的角色扮演对话奖励建模基准,涵盖从叙事管理到角色一致性和参与度的七种细粒度能力。在RoleRMBench上的评估揭示了通用奖励模型与人类判断之间存在显著且一致的差距,尤其是在叙事和风格维度。我们进一步提出了RoleRM,一种通过连续隐式偏好(CIP)训练的奖励模型,将主观评估重新表述为在多种结构化策略下的连续一致成对监督。全面实验表明,RoleRM在平均上超越了强大的开源和闭源奖励模型超过24%,在叙事连贯性和风格忠实度方面实现了显著提升。我们的发现突显了连续偏好表示和标注一致性的重要性,为人类中心对话系统中的主观对齐奠定了基础。
英文摘要:
Reward modeling has become a cornerstone of aligning large language models (LLMs) with human preferences. Yet, when extended to subjective and open-ended domains such as role play, existing reward models exhibit severe degradation, struggling to capture nuanced and persona-grounded human judgments. To address this gap, we introduce RoleRMBench, the first systematic benchmark for reward modeling in role-playing dialogue, covering seven fine-grained capabilities from narrative management to role consistency and engagement. Evaluation on RoleRMBench reveals large and consistent gaps between general-purpose reward models and human judgment, particularly in narrative and stylistic dimensions. We further propose RoleRM, a reward model trained with Continuous Implicit Preferences (CIP), which reformulates subjective evaluation as continuous consistent pairwise supervision under multiple structuring strategies. Comprehensive experiments show that RoleRM surpasses strong open- and closed-source reward models by over 24% on average, demonstrating substantial gains in narrative coherence and stylistic fidelity. Our findings highlight the importance of continuous preference representation and annotation consistency, establishing a foundation for subjective alignment in human-centered dialogue systems.