arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39882cs.CLcs.LG

LLM人格去学习

LLM Persona Unlearning

Kemou Li, Zhuan Shi, Qizhou Wang, Fengpeng Li, Negar Rostamzadeh, Golnoosh Farnadi, Jiantao Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出人格去学习问题,并构建基准PersonaUnlearnBench,发现标准方法失效,进而提出PaCE方法,通过对比目标与期望响应定位行为方向并训练状态转移,有效抑制目标人格,保持响应质量,实现权重级持久控制。

中文摘要 AI 辅助

预训练赋予大型语言模型(LLMs)与角色、风格、价值观和目标相关的广泛行为模式库。后训练教导条件性执行并使乐于助人的助手成为默认行为,但它并未从权重中抹除替代模式;因此,显式提示可以引出反复塑造判断、语言和行动的人格。在开放权重设置中,运行时控制可被移除,这促使了人格去学习:一种权重级编辑,使指定人格在未见上下文中难以被引出和执行。我们引入了PersonaUnlearnBench,一个针对模型特定的配对基准,涵盖来自三个家族的六种LLM和五种人格,具有对齐的遗忘/保留集、保留的指令改写和四轴评估。该基准表明,标准去学习方法无法可靠地抹除目标人格,而不牺牲有意义的生成或通用效用。因此,我们提出了PaCE,它比较目标响应和期望响应于相同问题,以定位内部行为方向,然后训练目标提示状态远离目标模式并朝向匹配的期望响应。实验表明,PaCE在中等效用成本下,一致地抑制目标人格,同时保持高响应质量和有用的对应行为。这些结果将人格去学习确立为一个独特的行为级编辑问题,以及一条通往对潜在LLM响应策略进行持久控制的实用途径。

英文摘要

Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-training teaches conditional enactment and makes a helpful Assistant the default, but it does not erase alternative modes from the weights; explicit prompts can therefore elicit personas that repeatedly shape judgment, language, and action. In open-weight settings, runtime controls can be removed, motivating persona unlearning: a weight-level edit that makes a designated persona difficult to elicit and enact on unseen contexts. We introduce PersonaUnlearnBench, a model-specific paired benchmark spanning six LLMs from three families and five personas, with aligned forget/retain sets, held-out instruction paraphrases, and four-axis evaluation. The benchmark shows that standard unlearning methods cannot reliably erase the target persona without sacrificing meaningful generation or general utility. We therefore propose PaCE, which compares target and desirable responses to the same questions to locate an internal behavior direction, then trains target-prompt states away from the target mode and toward the matched desirable response. Experiments show that PaCE consistently suppresses target personas with high response quality and useful counterpart behavior, at moderate utility cost. These results establish persona unlearning as a distinct behavior-level editing problem and a practical route toward persistent control of latent LLM response policies.

发表机构

  • State Key Laboratory of Internet of Things for Smart City, University of Macau(澳门大学智慧城市物联网国家重点实验室)
  • Mila – Québec AI Institute(米拉魁北克人工智能研究所)
  • McGill University(麦吉尔大学)
  • RIKEN AIP(日本理化学研究所先进智能研究中心)
  • King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
  • Google Research(谷歌研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑