发表机构
Emory University(埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通过大规模系统应用角色向量审计开放权重语言模型,编制特征清单并标记其性质,发现模型默认行为特点,引导在默认外特征效果好,且角色向量可探测行为组织。
AI 中文摘要
语言模型在训练后很大程度上决定了其行为表现,但仅通过提示无法揭示其具体表现、隐藏或抗拒的行为。角色向量作为激活空间中的行为方向可对此进行探究,但此前研究涉及的特征较少。本文首次大规模系统应用角色向量,编制了涵盖四个行为不同领域的53个特征清单,并标记了两个开放权重模型中各特征的性质。研究发现两个模型默认表现为有用的、面向任务的行为,引导在默认行为之外的特征上效果显著,且角色向量能探测行为组织。
英文摘要
What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits. We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural (expressed at baseline), steerable latent but amplifiable, or intractable (resistant to standard extraction). Both models default to helpful, task-oriented behavior: all nine agentic traits are natural, and their default clinician behavior matches a board-certified psychologist's independent desirability judgments on 16 of 17 traits. Steering produces its largest gains on traits these defaults exclude: hyperbole, hallucination, and sycophancy. The same asymmetry holds across all 171 generic-trait pairs: two steerable traits can collapse the composition, but pairs involving a default never do. Where standard extraction fails on a trait like "evil," a vector transferred from a fine-tuned variant still recovers it, with the residual refusals appearing inside the model's chain-of-thought. Persona vectors are most informative not as a set of controls but as a probe of behavioral organization.