发表机构
University of Melbourne; Hebrew University of Jerusalem; University of Tübingen(墨尔本大学; 耶路撒冷希伯来大学; 图宾根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过对比人类众包人格评分与LLM内部激活,发现模型特质表征结构与人类内隐人格结构高度一致,并建立了人机比较框架。
AI 中文摘要
一个世纪的心理学研究发现,人们用来描述他人的特质词汇各不相同,但这些特质之间的相互关系结构——哪些特质相互关联,哪些相互对立——在不同评价者和文化中却惊人地一致。我们测试了LLM(Qwen 2.5-7B-Instruct)是否在其内部特质表征中再现了这一结构。基于数百万条众包的对虚构角色的人格评分,我们构建了一个涵盖数百种特质的人类内隐人格矩阵;通过对比模型激活,我们构建了覆盖相同特质的匹配矩阵。这两种关系结构高度对齐(Mantel r = 0.77),且这种一致性在逐特质层面和整体层面均成立。模型特质表征的两个主导轴恢复了长期以来已知的组织人类人格印象的社会和智力维度,即社会温暖和智力能力。在保留的对话数据上,将模型激活投影到这些方向上所得到的人格轮廓与人类评分一致。这项工作建立了一个框架,能够实现对模型内部特质几何结构与人类人格印象共享结构之间全面且以人类为基准的比较。
英文摘要
A century of psychology has found that the trait words people use to describe one another vary, but the relational structure among those traits, which ones go together and which oppose, is strikingly consistent across raters and cultures. We test whether the LLM (Qwen 2.5-7B-Instruct) reproduces this structure in its internal trait representations. From millions of crowd-sourced personality ratings of fictional characters, we build a human implicit-personality matrix over hundreds of traits; from contrastive model activations, we build a matching matrix over the same traits. The two relational structures align strongly (Mantel r = 0.77), and the agreement holds trait by trait as well as in aggregate. Two dominant axes of the model's trait representations recover the social and intellectual dimensions long known to organize human personality impressions, social warmth and intellectual competence. On held-out dialogue, projecting model activations onto these directions yields personality profiles that agree with human ratings. This work establishes a framework that enables comprehensive, human-grounded comparison between internal model trait geometry and the shared structure of human personality impressions.
CommentsFindings of the Association for Computational Linguistics (EMNLP 2026)