arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12373cs.AI

面向LLM的鲁棒个性化对齐:缓解多轮对话中的角色漂移

Toward Robust Personalized Alignment for LLMs: Mitigating Persona Drift in Multi-Turn Dialogue

Youyuan Zhang, Siyuan Li, Fangming Liu, Jing Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对多轮对话中用户偏好变化导致的角色漂移问题,提出CORE方法分离回合级证据与持久状态修订,并引入PERSIST基准,在多个数据集上提升个性化对齐鲁棒性。

中文摘要 AI 辅助

角色漂移仍然是个性化语言模型面临的核心挑战,因为用户画像在长期交互中会不断演变,而非保持永久固定。因此,当用户偏好真正改变时,模型必须修订持久的角色状态,同时避免因短暂、模糊或未解决的观察而触发更新。我们提出CORE,它将回合级证据与持久角色状态修订相分离,并通过基于不确定性的信念修订来选择性更新有根据的用户偏好。我们还引入了PERSIST,一个用于在序列交互压力下评估角色状态鲁棒性的留出后锚定基准,涵盖模糊性、冲突和受控社会影响。在ALOE、PersonaChat和PERSIST上,CORE提升了个性化对齐和鲁棒性,并在归一化闭槽状态保真度上取得了互补性增益。人工评估和机制控制进一步支持显式更新控制优于仅增强生成或持久记忆。

英文摘要

Persona drift remains a central challenge for personalized language models, as user profiles evolve over long interactions rather than remain permanently fixed. Models must therefore revise persistent persona states when preferences genuinely change, while avoiding updates driven by transient, ambiguous, or unresolved observations. We propose CORE, which separates turn-local evidence from persistent persona-state revision and selectively updates grounded user preferences through uncertainty-aware belief revision. We also introduce PERSIST, a held-out post-anchor benchmark for persona-state robustness under sequential interaction stress, covering ambiguity, conflict, and controlled social influence. Across ALOE, PersonaChat, and PERSIST, CORE improves personalized alignment and robustness, with complementary gains in normalized closed-slot state fidelity. Human evaluation and mechanistic controls further support explicit update control beyond stronger generation or persistent memory alone.

发表机构

  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
  • Peng Cheng Laboratory(鹏城实验室)
  • Huazhong University of Science and Technology(华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑