arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PHASE-Tree:建模长回合角色扮演对话中的角色状态演化

PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue

Bo Tang, Jianan Yang, Junyi Zhu, Yiquan Wu, Rui Zhao, Zhengyu Yang, Yang Zhang, Feiyu Xiong, Zhiyu Li, Jiajun Shen

arXiv 2608.06975首次发表:更新:

发表机构

MemTensor (Shanghai) Technology; KU Leuven; Zhejiang University; University of Chinese Academy of Sciences; Sinar Mas Paper (China) Investment Co., Ltd; The Hong Kong Polytechnic University(MemTensor(上海)科技; 鲁汶大学; 浙江大学; 中国科学院大学; 金光纸业(中国)投资有限公司; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长回合角色扮演对话的角色状态演化问题,提出多时间尺度角色状态树模型PHASE-Tree,构建基准LongEvoRoleBench并验证其在角色生成任务上的性能优势。

AI 中文摘要

长回合角色扮演要求角色随叙事演化时仍保持可识别性,但现有工作存在两方面不足:表示通常为静态画像,无法在不破坏未变更特质的情况下进行局部更新;基准主要测试角色人格保留与记忆召回,而非模型是否基于角色当前演化状态发言。我们针对这两点提出解决方案:PHASE-Tree是一种多时间尺度的角色状态树,包含不可变的身份根节点与可变的人格、会话、时刻层,使每个可变字段成为可寻址目标,用于在单回合内及跨回合的局部更新;该模型通过显式文本提供或隐式参数适配来条件化生成。为衡量演化状态下的生成效果,我们引入LongEvoRoleBench,其将4个长对话语料库(用于跨回合演化)与4个短对话语料库(用于场景内状态跟踪检查)配对,采用统一的下一轮发言协议。在长对话核心任务上,文本版PHASE-Tree在12个数据集-指标单元中,对比内部变体有11个单元排名第一,对比所有外部文本基线则12个单元均排名第一,分别将角色级、语义级和嵌入级分数提升19.7%、12.4%和15.1%。在一项200条回复的盲测研究中,人类评分与GPT-4.1评判者的皮尔逊相关系数为0.65;在描述性的n=10个PT和NR提示子集上,总体差异为+0.20;长对话的语义优势在大语言模型评判者与生成主干中均持续存在。

英文摘要

Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on two fronts: representations are typically static profiles that cannot be updated locally without destabilizing unchanged traits, and benchmarks mainly test persona preservation and memory recall rather than whether a model speaks from a character's currently evolved state. We address both. PHASE-Tree is a multi-timescale character-state tree with an immutable identity root and mutable persona, session, and moment layers, making each mutable field an addressable target for localized within- and cross-episode updates. It conditions generation through explicit textual provision or implicit parametric adaptation. To measure evolved-state generation, we introduce LongEvoRoleBench, which pairs four long-dialogue corpora for cross-episode evolution with four short-dialogue corpora as within-scene state-tracking checks, under a unified next-utterance protocol. On the long-dialogue core, textual PHASE-Tree ranks first in 11 of 12 dataset-metric cells against internal variants and all 12 cells against external textual baselines, improving character-level, semantic, and embedding scores by 19.7%, 12.4%, and 15.1% respectively. In a blinded 200-response study, human ratings correlate with the GPT-4.1 judge (Pearson r= 0.65); on descriptive n= 10 PT and NR prompt subsets, the Overall difference is +0.20. The long-dialogue Sem advantage persists across LLM judges and generation backbones.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑