AI 中文总结
本研究基于Funder人格三元框架,提出用于发现、控制和验证大语言模型类特质表征的框架,证实LLMs存在可连接内部状态、情境与行为的可控类特质表征。
AI 中文摘要
人类人格理论将特质描述为并非由单一分数捕捉的孤立属性,而是通过人、情境与行为三者的相互作用所表现出的稳定个体倾向。现有针对大语言模型(LLMs)人格相关行为的研究主要聚焦于人格条件下产生的输出,刻画可观察的特质相关表达,但缺乏内部人格相关表征存在的机制证据、其跨情境的表现方式,以及这些表征如何塑造特定行为。基于Funder的人格三元框架,我们将其三个组成部分适配到LLM分析中:“人”指与人格相关的内部表征,“情境”指能提供特质相关反应的上下文,“行为”指更广泛社会任务上的反应模式。我们提出了一个用于发现、控制和验证LLMs中类特质表征的框架:首先,利用基于共享情境的对比行为对,通过稀疏自动编码器(SAE)分解识别与人格特质对立两极相关的稀疏内部特征,通过对情境的行为影响、token级激活模式及对改述的鲁棒性验证其特质相关性;其次,特征级干预在一组独立的多样化情境中诱导双向特质相关的转变,同时保持反应有效性,证明其在不同情境下的一致表现;最后,将相同干预应用于社会智能任务,揭示出与人类人格研究结果一致的利弊权衡模式的行为变化,提供了超越人格分数的行为层面验证。我们的发现证明LLMs包含可控制的类特质表征,能连接内部状态、情境表现与行为结果。
英文摘要
Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait-related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality-related internal representations, Situation as contexts that afford trait-relevant responses, and Behavior as response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.