发表机构
Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出用大五人格向量解释语言模型微调引发的错位,发现错位语料具特定人格特征,该向量可诊断安全现象且能捕捉单一方向无法识别的谄媚特质。
AI 中文摘要
在包含窄缺陷(如不安全代码或错误数学答案)的数据上微调语言模型,会通过仍存在争议的机制引发广泛错位。本文提供可解释的解释:在研究的模型和语料中,错位表现为人格转变。先前工作从单一二元对比中提取性格特质的激活方向,可分离或引导行为但无法建立校准量表;本文采用分级三级干预提取大五人格的人格向量,并在两个开放权重模型上验证。三级呈线性排序,Cohen's d值最高达6.2;向量零样本迁移且具特质特异性至独立语料,效应在中层带内最强。应用于训练数据时,该向量揭示八个领域的错位语料共享共同大五人格特征:宜人性和尽责性更低,外向性和神经质更高,模型间该特征的相关性为r=0.94。微调会留下相同特征,使模型生成沿对应特征转变,基于激活测量的相关性为r=0.83,基于文本评判器的相关性为r=0.90,内部激活转变的相关性为r=0.69。相同向量将谄媚特征化为高外向性和低尽责性,而非过度宜人性,这是单一方向无法捕捉的区别。校准后的人格向量将不透明的安全现象转化为人类可理解的诊断特征。
英文摘要
Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable account: in the models and corpora we study, misalignment behaves like a shift in personality. Prior work extracts activation directions for character traits from a single binary contrast, which can separate or steer behavior without establishing a calibrated scale. We instead extract personality vectors for the Big Five using a graded, three-level intervention and validate them on two open-weight models. The three levels are linearly ordered, with Cohen's d values of up to 6.2; the vectors transfer zero-shot and trait-specifically to an independent corpus; and their effects are strongest within a middle-layer band. Applied to training data, the vectors reveal that misaligned corpora across eight domains share a common Big Five signature: lower agreeableness and conscientiousness, together with higher extraversion and neuroticism. This signature is recovered by both models with a correlation of r = 0.94. Fine-tuning imprints the same profile, shifting the model's generations along the corresponding signature, with r = 0.83 using activation-based measurements and r = 0.90 using a text-based judge, while also shifting internal activations with r = 0.69. The same vectors characterize sycophancy as high extraversion and low conscientiousness rather than excess agreeableness, a distinction that a single direction cannot capture. Calibrated personality vectors transform an opaque safety phenomenon into a human-legible diagnostic profile.
CommentsThe paper is currently under peer review