发表机构
Northeastern University; Stanford University; Google DeepMind(东北大学; 斯坦福大学; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出人格层级模型解释大语言模型微调中的上下文泛化现象,并通过120个模型实验验证,同时提出人格保留正则化方法有效抑制不期望的泛化,降低奖励黑客行为。
AI 中文摘要
语言模型通常在固定上下文(如通用系统提示、人格或领域特定指令)下进行微调,然而学习到的行为有时局限于该上下文,有时则广泛泛化到未见过的上下文。我们提出人格层级模型来解释这一现象:一个共享的默认人格会影响跨上下文的行为。在该模型下,修改共享人格的微调能促进更广泛的迁移,而对局部人格的修改则更局限于特定上下文。在涵盖四种行为和15种训练上下文的120个微调模型中,泛化狭窄性与训练上下文人格和默认人格之间的相似性呈正相关(Qwen3-4B的Pearson相关系数r=0.72)。在默认上下文下的先前微调可以拓宽后续在其他上下文训练中的泛化。将上下文响应与默认人格响应对齐会产生更强的效果。最后,我们提出人格保留正则化(PPR)来限制不期望的上下文泛化。在强化学习中,PPR将奖励黑客行为从42-55%降低到每个评估提示下至多0.2%,同时保持准确率提升。这些结果支持人格层级模型作为上下文泛化的一种解释,并可能激励未来对意外泛化的控制,以更好地对齐大语言模型。
英文摘要
Language models are routinely fine-tuned under a fixed context, such as a generic system prompt, persona or domain-specific instruction, yet the learned behavior sometimes stays confined to that context and sometimes broadly generalizes to unseen contexts. We propose the Persona Hierarchy Model to explain this: a shared default persona influences behavior across contexts. Under this model, fine-tuning that modifies the shared persona promotes broader transfer, whereas changes to local personas remain more context-specific. Across 120 fine-tuned models spanning four behaviors and 15 training contexts, generalization narrowness positively correlates with the similarity between the training context's persona and the default persona (Pearson's r = 0.72 for Qwen3-4B). Prior fine-tuning under the default context can broaden generalization in subsequent training under other contexts. Aligning contextual responses with default-persona responses produces stronger effects. Finally, we propose persona-preserving regularization (PPR) to confine undesired contextual generalization. In RL, PPR cuts reward hacking from 42-55% to at most 0.2% under every evaluated prompt while retaining accuracy gains. These results support the Persona Hierarchy Model as an explanation for contextual generalization and can motivate future controls on unintended generalization for better alignment of LLMs.