arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OSPD:用于人格一致对话的在线策略自蒸馏

OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue

Rui Xu, Yikai Zhang, Aili Chen, Zicheng Zhao, Xu Yinghui, Libo Wu

arXiv 2609.34418首次发表:更新:

发表机构

Fudan University; Shanghai Innovation Institute(复旦大学; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出在线策略自蒸馏框架OSPD,通过不对称信息下的自蒸馏、角色感知散度切换及渐进特征掩蔽课程,在不依赖外部教师或奖励模型的情况下,显著提升多轮对话中的人格一致性。

AI 中文摘要

在多轮对话中保持人格一致性,对于角色扮演语言模型而言仍是一个核心挑战。来自外部教师模型的离线策略蒸馏会导致分布不匹配,且这种不匹配会在对话轮次间累积;而强化学习则面临主观人格保真度中固有的奖励模糊性问题。我们提出了OSPD,一种在线策略自蒸馏框架,其中同一模型在不对称信息下同时充当教师和学生:教师接收完整的角色档案,而学生仅看到简要摘要,且学生从其自身策略生成轨迹。我们发现,角色扮演对话中的教师置信度呈现双峰结构——在角色关键令牌处急剧峰值,而在通用话语处则弥散——并引入角色感知的散度切换以匹配该结构。渐进式特征掩蔽课程进一步促使角色知识沿语义维度进行分阶段内化。在CharacterBench、CharacterEval和SocialBench上的实验表明,OSPD在人格一致性方面显著优于监督微调和多轮强化学习基线,且无需任何外部教师模型或奖励模型。

英文摘要

Maintaining persona consistency across multi-turn dialogues remains a core challenge for role-playing language models. Off-policy distillation from external teachers incurs distribution mismatch that compounds across dialogue turns, while reinforcement learning struggles with reward ambiguity inherent in subjective persona fidelity. We propose OSPD, an on-policy self-distillation framework where the same model serves as both teacher and student under asymmetric information: the teacher receives a complete character profile while the student sees only a brief summary, and the student generates trajectories from its own policy. We find that teacher confidence in role-playing dialogue exhibits a bimodal structure---sharply peaked at character-critical tokens yet diffuse at generic utterances---and introduce role-aware divergence switching to match this structure. A progressive trait masking curriculum further forces staged internalization of character knowledge along semantic dimensions. Experiments on CharacterBench, CharacterEval, and SocialBench show that OSPD substantially improves persona consistency over supervised fine-tuning and multi-turn RL baselines, without requiring any external teacher or reward model.

CommentsAccepted to EMNLP 2026. 22 pages, including references and appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑