arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KuaiRP系列角色扮演模型技术报告

KuaiRP Series Role-playing Models Technical Report

Yipeng Wang, Ziwei Zhang, Jiahui Zhang, Qi Gan, Kai Sheng

arXiv 2609.11127首次发表:更新:

发表机构

Kuaishou GameMind Lab(快手游戏心智实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KuaiRP系列角色扮演模型通过多阶段训练流程,包括标准化模板SFT、规则复合奖励RL和两阶段同策略蒸馏,平衡深度领域知识注入与通用智能体能力保持,实现高保真角色扮演和低部署成本。

AI 中文摘要

本文介绍了KuaiRP系列角色扮演模型的完整技术方案。我们旨在为专用角色扮演模型实现四个核心目标:简化的提示工程、高度稳定的输出质量、内置的领域世界知识以及小参数规模下的高效部署。然而,有效注入深度领域知识往往会导致模型通用智能体能力的严重灾难性遗忘。为克服这一权衡,我们提出了一种多阶段训练流程。首先,我们设计了一个标准化的角色模板,并基于用户行为模拟和反向画像过滤构建了SFT数据流水线。接下来,我们在强化学习(RL)阶段利用基于规则的复合奖励函数来消除长度膨胀和重复生成等常见退化现象。最后,为恢复在SFT和RL阶段受损的通用能力,我们提出了一种新颖的自蒸馏范式,即使用配备累积散度衰减(CDD)的两阶段同策略蒸馏(OPD)。通过将领域自适应模型作为教师模型,原始基础模型作为学生模型,我们有效平衡了深度领域知识注入与通用智能体能力的保持。实验结果表明,KuaiRP模型不仅在我们目标领域内的角色扮演保真度上匹配当前最先进的专有模型,而且成功恢复了通用智能体能力,同时保持了极低的部署成本。

英文摘要

This paper introduces the complete technical solution for the KuaiRP series of role-playing models. We aim to achieve four core objectives for a dedicated role-playing model: simplified prompt engineering, highly stable output quality, built-in domain world knowledge, and high-efficiency deployment with a small parameter size. However, effectively injecting deep domain knowledge often leads to a severe catastrophic forgetting of the model's general agent capabilities. To overcome this trade-off, we propose a multi-stage training pipeline. First, we design a standardized character template and construct an SFT data pipeline based on user behavior simulation and reverse profile filtering. Next, we utilize a rule-based composite reward function during the Reinforcement Learning (RL) phase to eliminate common degradation phenomena like length expansion and repetitive generation. Finally, to recover the general capabilities compromised during SFT and RL, we propose a novel self-distillation paradigm using Two-stage On-Policy Distillation (OPD) equipped with Cumulative-Divergence Decay (CDD). By using the domain-adapted model as the teacher and the original base model as the student, we effectively balance deep domain knowledge injection with the preservation of general agent capabilities. Experimental results demonstrate that the KuaiRP models not only match the current state-of-the-art proprietary models in role-playing fidelity within our target domains, but also successfully recover general agent capabilities, maintaining extremely low deployment costs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑