你的大语言模型,你的风格:用于大语言模型行为控制的行为模式轴
Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control
浏览论文内容
中文总结 AI 辅助
本研究提出情境化行为数据框架,构建3200个对比性行为场景,发现LLMs有稳定且模型特异性的行为特征,提出行为模式轴控制LLM行为,表明其类人格倾向是可测量可控的行为模式。
中文摘要 AI 辅助
大语言模型(LLMs)越来越多地在交互场景中运行,其行为风格会影响用户体验、安全性和下游决策。现有的LLM人格研究大多依赖在第一人称场景中进行的自我报告问卷,使得生成的特征对表面诱导选择敏感,且难以与具体的模型行为建立关联。在本研究中,我们引入了一种情境化行为数据(B-data)框架,用于研究和控制LLM的行为人格。我们构建了3200个对比性行为场景,涵盖20种行为模式和4种提示语域,这些场景基于经过验证的心理测量维度,如大五人格量表第二版(BFI-2)、风险决策偏好量表(DOSPERT)和六因素人格模型(HEXACO)。利用该框架,我们发现LLMs表现出稳定且具有模型特异性的行为特征,同时也揭示了在第一人称决策、提供建议和任务执行中存在的语域依赖型转变。随后,我们证明这些行为模式可以通过行为模式轴(BMAs)进行控制,BMAs是从对比性行为轨迹中推导出来的激活空间方向。与更容易出现特质漂移的响应衍生BMAs相比,思维衍生BMAs能更忠实地捕捉预期的行为机制,并对情境化行为风格提供更精准的控制。我们的结果表明,LLM类人格倾向更适合被理解为可测量和可控的行为模式,而非基于具体交互情境的抽象自我报告特质。我们的代码和数据可在该https链接获取。
英文摘要
Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report questionnaires administered in first-person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in concrete model behavior. In this work, we introduce a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality. We construct 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric facets such as BFI-2, DOSPERT, and HEXACO. Using this framework, we find that LLMs exhibit stable and model-specific behavioral profiles, while also revealing register-dependent shifts across first-person decisions, advice-giving, and task execution. We then show that these behavioral patterns can be controlled through Behavioral Mode Axes (BMAs), activation-space directions derived from contrastive behavioral traces. Compared with response-derived BMAs, which are more prone to trait drift, thought-derived BMAs more faithfully capture the intended behavioral mechanism and provide cleaner control over situated behavioral styles. Our results suggest that LLM personality-like tendencies are better understood not as abstract self-report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts. Our code and data are available at https://github.com/lhz191/LLM-Behavioral-Personality.