人格的几何学:用荣格认知功能进行激活引导
The Geometry of Personality: Activation Steering with Jungian Cognitive Functions
浏览论文内容
中文总结 AI 辅助
研究能否用荣格认知功能将人格表示为认知过程并控制,引入含评估协议和数据集的框架,通过实验表明可有效控制,还揭示人格信息所在层、引导向量关系等,为大语言模型人格表示提供新见解并建立研究框架。
中文摘要 AI 辅助
激活引导能够实现对大语言模型的控制与解释,但现有工作主要通过诸如大五人格等静态特质框架来塑造人格。我们研究能否用八种荣格认知功能将人格表示并控制为一组认知过程。为此,我们引入了一个包含荣格评估协议和超2100个角色扮演人物叙述数据集的框架。对Llama - 3.1 - 8B进行的激活引导向量提取和评估实验表明,通过激活引导可对所有八种认知功能进行有效的单调控制。除了可控性,我们的分析还揭示:人格信息集中在中间的Transformer层;引导向量呈现出与理性和非理性功能区别一致的结构化几何关系;有效的多维引导方向无法作为单功能方向的线性组合恢复。这些发现为大语言模型激活空间中人格的表示提供了新见解,并建立了一个研究可解释、有效且多维人格控制的框架。
英文摘要
Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled as a set of cognitive processes using the eight Jungian Cognitive Functions. To this end, we introduce a framework comprising a Jungian evaluation protocol and a dataset of over 2,100 role-playing character narrations. Activation steering vector extraction and evaluation experiments on Llama-3.1-8B demonstrate effective monotonic control over all eight cognitive functions through activation steering. Beyond controllability, our analysis reveals that: 1. personality information is concentrated in middle transformer layers; 2. steering vectors exhibit structured geometric relationships consistent with distinctions between rational and irrational functions; 3. effective multi-dimensional steering directions cannot be recovered as linear combinations of single-function directions. These findings provide new insights into the representation of personality in LLM activation space and establish a framework for studying interpretable, effective, and multi-dimensional personality control.
发表机构
- University of Glasgow(格拉斯哥大学)
- Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。