R²A:通过角色表征学习与运行时对齐学习角色策略
R$^2$A: Learning Persona Policies Through Persona Representation Learning and Runtime Alignment
- College of AI, Tsinghua University(清华大学人工智能学院)
- Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对静态角色触发表现不一致的问题,提出R²A方法,通过角色表征学习与运行时对齐学习角色策略,在12种评估设置中优于基础模型与静态角色触发。
AI中文摘要:
同一角色(Persona)的行为在一种情境中可能有益,而在另一种情境中则有害,这导致静态角色触发在不同任务中表现不一致。我们提出角色选择-实现框架(Persona Selection--Realization Framework),该框架通过潜在角色状态对行为生成进行建模,并将其分解为角色选择(Persona Selection)与角色实现(Persona Realization)两个部分。静态角色触发与理想角色策略在这两个部分的差异分别定义为选择差距(Selection Gap)与实现差距(Realization Gap)。基于该框架,我们提出R²A,这是一种学习角色策略的两阶段方法。角色表征学习(Persona Representation Learning)使用结构化的“谁-如何-什么”(Who--How--What)表示,对目标角色的目标、条件行为原则及轨迹级表现进行编码。随后的角色运行时对齐(Persona Runtime Alignment)会移除显式角色规范,并利用任务反馈共同校准行为选择与轨迹实现。在覆盖本研究中“可问责专业角色(Accountable-Professional Persona)”四项原则的12种评估设置中,R²A整体表现优于基础模型与静态角色触发。消融实验结果进一步表明,角色表征学习对于防止运行时对齐产生行为不平衡的策略,以及实现更稳定的角色策略学习至关重要。
英文摘要:
The same Persona behavior can be beneficial in one context but harmful in another, causing static Persona elicitation to perform inconsistently across tasks. We introduce the Persona Selection--Realization Framework, which models behavior generation through a latent Persona state and decomposes it into Persona Selection and Persona Realization. The discrepancies between static Persona elicitation and an ideal Persona policy in these two components define the Selection Gap and Realization Gap, respectively. Building on this framework, we propose R$^2$A, a two-stage approach for learning Persona policies. Persona Representation Learning uses structured Who--How--What presentations to encode the target Persona's objective, conditional behavioral principles, and trajectory-level manifestations. Persona Runtime Alignment then removes the explicit Persona specification and jointly calibrates behavior selection and trajectory realization using task feedback. Across 12 evaluation settings covering the four principles of the Accountable-Professional Persona studied in this work, R$^2$A overall outperforms both the base model and static Persona elicitation. Ablation results further show that Persona Representation Learning is critical for preventing Runtime Alignment from producing behaviorally imbalanced policies and for achieving more stable Persona policy learning.