发表机构
The University of Tokyo; RIKEN Center for Advanced Intelligence Project(东京大学; 理化学研究所先进智能项目中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SPIN,一种基于认知-情感人格系统的推断流水线,通过三次零样本LLM调用分离稳定人格与情境状态,在多个基础模型上显著提升社会心理行为模拟的对齐性能。
AI 中文摘要
大语言模型越来越多地被用于在社会科学和行为研究中模拟人类参与者,然而静态人格提示通常直接将参与者档案和实验场景映射到响应,从而将稳定的倾向与情境特定的解释纠缠在一起。为解决这一局限性,我们引入了SPIN,一种受认知-情感人格系统启发的推断流水线,用于模拟人类社会心理行为。具体而言,SPIN通过三次零样本大语言模型调用实现这一结构化推断过程:编译一个任务无关的参与者核心,引发特定条件下的认知-情感状态,并从这些状态中读出决策,从而在重用稳定人格结构的同时,通过显式的状态表示路由每个试验特定的响应。我们在两个重建的社会心理研究系列(涵盖不确定性推理和多元无知)上,跨四个基础大语言模型评估了SPIN。与空白、人口统计、叙事和思维链提示变体相比,SPIN在基础大语言模型和研究系列中始终提供最强的整体对齐性能。消融研究和状态分析进一步表明,人格编译和结构化状态引发都有助于性能提升,并且引发的状态在信息性和规范性条件下可解释地变化。这些结果表明,结构化的人格-状态推断可以在基准层面提升行为对齐,超越更丰富的人格描述或通用的多步推理。
英文摘要
Large language models are increasingly used to simulate human participants in social and behavioral studies, yet static persona prompting typically maps a participant profile and an experimental scenario directly to a response, entangling stable dispositions with situation-specific interpretations. To address this limitation, we introduce \textbf{SPIN}, a cognitive-affective personality system-inspired inference pipeline for simulating human social-psychological behavior. Specifically, SPIN implements this structured inference process through three zero-shot LLM calls that compile a task-blind participant core, elicit condition-specific cognitive-affective states, and read out decisions from those states, thereby reusing stable personality structure while routing each trial-specific response through an explicit state representation. We evaluate SPIN on two reconstructed social-psychological study families spanning uncertainty reasoning and pluralistic ignorance, across four base LLMs. Compared with blank, demographic, narrative, and chain-of-thought prompt variants, SPIN consistently delivers the strongest overall alignment performance across base LLMs and study families. Ablations and state analyses further show that both personality compilation and structured state elicitation contribute to the gains, and that the elicited states shift interpretably across informational and normative conditions. These results suggest that structured personality-state inference can improve benchmark-level behavioral alignment beyond richer persona descriptions or generic multi-step reasoning.
CommentsAccepted at NeurIPS 2026