用于社会模拟的语言模型角色引导
Role Steering of Language Models for Social Simulations
- Georgia Institute of Technology(佐治亚理工学院)
- ML Alignment & Theory Scholars (MATS)(ML对齐与理论学者组织(MATS))
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- University of Oxford(牛津大学)
- Google DeepMind(谷歌DeepMind)
- OpenAI(开放人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出针对语言模型智能体的角色引导筛选工作流,在OLMo-3-7B-Instruct上验证其角色画像对齐度更高且词汇多样性更好,发现需按角色选引导系数而非统一高强度设置。
AI中文摘要:
由语言模型智能体构建的社会模拟需要可在智能体被置于模拟群体前检查的角色条件行为。我们提出一种针对角色条件智能体的激活引导筛选工作流:定义角色画像、提取角色特定方向、遍历四个引导系数、评估角色画像对齐度,对每个候选配置通过或标记。在OLMo-3-7B-Instruct上,我们将该工作流应用于包含275个混合角色的清单,搭配228个与角色无关的问题、经GPT-4.1-mini提示生成的角色参考,以及GPT-4.1-mini评判者。角色特定方向的评判角色画像对齐度高于现有角色向量研究中的助手轴方向控制,在测试网格上的平均总得分分别为63.2和41.1;同时其保留了高词汇多样性,而控制方法在系数增大时词汇多样性急剧下降。角色级筛选是主要实用产出:多数角色随引导增强而表现提升,但有38个角色在全部六个测量维度上表现下降,这表明模拟构建者应针对每个角色选择系数,而非部署统一的高强度设置。我们将代码和评估工件公开于此httpsURL。
英文摘要:
Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated population. We introduce an activation-steering screening workflow for role-conditioned agents: define a role profile, extract a role-specific direction, sweep four steering coefficients, evaluate role-profile alignment, and pass or flag each candidate configuration. On OLMo-3-7B-Instruct, we apply the workflow to a mixed 275-role inventory with 228 role-agnostic questions, GPT-4.1-mini prompted role references, and GPT-4.1-mini judges. Role-specific directions receive higher judged role-profile alignment than an assistant-axis directional control from prior persona-vector work, with mean overall scores of 63.2 versus 41.1 across the tested grid. They also preserve high lexical diversity, while the control drops sharply at larger coefficients. The role-level screen is the main practical output: most roles improve as steering increases, but 38 roles decline across all six measured dimensions, showing why simulation builders should choose coefficients per role rather than deploy a uniform high-strength setting. We make our code and evaluation artifacts available at https://anonymous.4open.science/r/anonymous-research-code-5F03/.