AI 中文总结
通过心理测量学评估大语言模型的大五人格测试,发现其不适用于LLMs,无法捕捉模型间差异且因子结构失效,建议开发针对LLMs的评估框架。
AI 中文摘要
人类人格量表越来越多地被用于描述大型语言模型(LLMs)、比较系统以及为下游治理主张提供信息。然而,这些量表是为人类开发和验证的,尚不清楚它们是否适用于LLMs。我们对LLMs中的大五人格测量进行了系统的心理测量学评估。我们提出三个研究问题:大五人格量表是否a)恰当地描述LLMs,b)捕捉模型间的个体差异,以及c)反映与人类人格一致的内部因素。我们评估了五个候选大五人格量表的内容效度,并将获胜量表施测于N=244个不同模型,涵盖49个模型家族。首先,我们发现为LLMs改编的大五人格条目可以达到足够的内容效度,而原始的人类开发条目则不能。其次,大五人格量表未能捕捉LLMs之间的有意义差异:我们发现模型间变异性低,仅占总得分方差的3%。第三,LLMs的响应未能恢复大五人格的五因子结构,其中四个大五维度合并为一个(r>=0.92)。基础模型与指令调优模型变体的直接比较表明,对齐训练系统地使大五人格得分向社会期望特质偏移。这些发现表明,大五人格得分在LLMs中并未测量出与人类人格等价的结构。将人类人格框架应用于LLMs会产生误导性的特征描述,这些描述被用于对LLMs进行基准测试、比较和治理。我们强调需要开发针对LLMs的评估框架,而不是未经验证地采用人类构念。
英文摘要
Human personality inventories are increasingly used to characterize large language models (LLMs), compare systems, and inform downstream governance claims. Yet, these inventories were developed and validated for humans, and it remains unclear whether they are valid for non-human systems. We present a systematic psychometric evaluation of Big Five personality measurement in LLMs. We ask three research questions: Do Big Five inventories a) appropriately describe LLMs, b) capture meaningful differences between models, and c) reflect internal factors consistent with human personality? We assess the content validity of five candidate Big Five inventories and administer the best-performing inventory to N = 264 LLMs spanning 50 model families. Our findings are threefold. First, Big Five items adapted for LLMs achieve acceptable content validity, whereas the original human-developed items do not. Second, Big Five inventories fail to capture meaningful differences across LLMs: between-model variance accounts for only 7% - 17% of the total score variance. Third, LLMs responses do not reproduce the canonical Big Five five-factor structure of human personality, with four of the five personality facets collapsing into one (r >= .90). Moreover, comparisons between base and instruction-tuned variants suggest that alignment training shifts Big Five scores toward socially desirable profiles. These findings demonstrate that Big Five inventories do not measure a construct equivalent to human personality in LLMs. Thus, using human personality frameworks to characterize, benchmark, compare, or govern LLMs risks producing misleading conclusions. We highlight the need for evaluation frameworks that are specifically designed and validated for LLMs, rather than transferring human psychological constructs without first establishing their validity.
Comments12 pages, 3 tables, 4 figures; Accepted for publication at the Ninth AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026)