发表机构
AI Institute; Harvard Business School(人工智能研究所; 哈佛商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
为计算复制品引入心理测量基准,通过固定身份人格经历事件序列,从内部效度和外部效度两维度评分,测试36个LLM,发现多数能恢复关系方向和方差划分,但个体内部结构难恢复,中等规模模型Gemma-3-27B表现最佳。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被部署用于模拟人类行为,充当人类受试者的计算复制品。然而,人类真实的心理体验难以进行基准测试,尤其是当其随时间展开时。我们为计算复制品引入了一个心理测量基准:即通过一系列不断演变的事件携带固定身份的人格(personas)。该基准完全基于已发表的常模和元分析效应构建,对心理现实主义的两个维度进行评分。第一个维度是内部效度,用于量化生成的轨迹是否重现了重复人类测量的内部结构:分布、人与人之间与个体内部方差划分、时间依赖性和范围。第二个维度是外部效度,用于量化复制品是否恢复了已确立的特质、状态和指标关系。在来自九个开发者的36个开放权重和专有LLM(1B-671B参数)中,大多数恢复了已确立关系的方向(平均一致性84.3%)和方差划分(34个中的26个),然而个体内部相关性和分布结构却难以恢复。内部效度与规模和能力无关:一个中等规模的开放LLM(Gemma-3-27B)在这两个维度之间取得了最佳权衡。该基准是在跨领域(如市场营销、医疗保健)进行因果推断时使用计算复制品的先决条件,并将个体内部基础识别为未来的核心挑战。
英文摘要
Large language models (LLMs) are increasingly deployed to simulate human behavior, acting as computational replicas of human subjects. Yet the lived psychological experience of humans is difficult to benchmark, particularly as it unfolds over time. We introduce a psychometric benchmark for computational replicas: personas that carry a fixed identity through an evolving sequence of events. Built entirely from published norms and meta-analytic effects, the benchmark scores two dimensions of psychological realism. The first, internal validity, quantifies whether generated trajectories reproduce the internal structure of repeated human measurement: distributions, the between- versus within-person variance partition, temporal dependence, and range. The second, external validity, quantifies whether replicas recover established trait, state, and indicator relations. Across 36 open-weight and proprietary LLMs from nine developers (1B-671B parameters), most recover the direction of established relations (84.3% mean agreement) and the variance partition (26 of 34), yet the within-person correlation and distributional structure elude recovery. Internal validity is independent of scale and capability: a mid-size open LLM (Gemma-3-27B) strikes the best trade-off between the two dimensions. The benchmark is a precondition for using computational replicas in causal inference across domains (e.g., marketing, healthcare), and identifies within-person grounding as the central challenge ahead.