arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10703cs.LGcs.AIcs.CLcs.HC

你的大语言模型,你的风格:用于大语言模型行为控制的行为模式轴

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出情境化行为数据框架,构建3200个对比性行为场景,发现LLMs有稳定且模型特异性的行为特征,提出行为模式轴控制LLM行为,表明其类人格倾向是可测量可控的行为模式。

中文摘要 AI 辅助

大语言模型(LLMs)越来越多地在交互场景中运行,其行为风格会影响用户体验、安全性和下游决策。现有的LLM人格研究大多依赖在第一人称场景中进行的自我报告问卷,使得生成的特征对表面诱导选择敏感,且难以与具体的模型行为建立关联。在本研究中,我们引入了一种情境化行为数据(B-data)框架,用于研究和控制LLM的行为人格。我们构建了3200个对比性行为场景,涵盖20种行为模式和4种提示语域,这些场景基于经过验证的心理测量维度,如大五人格量表第二版(BFI-2)、风险决策偏好量表(DOSPERT)和六因素人格模型(HEXACO)。利用该框架,我们发现LLMs表现出稳定且具有模型特异性的行为特征,同时也揭示了在第一人称决策、提供建议和任务执行中存在的语域依赖型转变。随后,我们证明这些行为模式可以通过行为模式轴(BMAs)进行控制,BMAs是从对比性行为轨迹中推导出来的激活空间方向。与更容易出现特质漂移的响应衍生BMAs相比,思维衍生BMAs能更忠实地捕捉预期的行为机制,并对情境化行为风格提供更精准的控制。我们的结果表明,LLM类人格倾向更适合被理解为可测量和可控的行为模式,而非基于具体交互情境的抽象自我报告特质。我们的代码和数据可在该https链接获取。

英文摘要

Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report questionnaires administered in first-person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in concrete model behavior. In this work, we introduce a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality. We construct 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric facets such as BFI-2, DOSPERT, and HEXACO. Using this framework, we find that LLMs exhibit stable and model-specific behavioral profiles, while also revealing register-dependent shifts across first-person decisions, advice-giving, and task execution. We then show that these behavioral patterns can be controlled through Behavioral Mode Axes (BMAs), activation-space directions derived from contrastive behavioral traces. Compared with response-derived BMAs, which are more prone to trait drift, thought-derived BMAs more faithfully capture the intended behavioral mechanism and provide cleaner control over situated behavioral styles. Our results suggest that LLM personality-like tendencies are better understood not as abstract self-report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts. Our code and data are available at https://github.com/lhz191/LLM-Behavioral-Personality.

补充信息

↑