大型语言模型中自我报告的原型与行为失败
Self-reported archetypes and behavioral failures in Large Language Models
AI总结:
本研究通过自我报告映射22个LLM的人格原型,发现闭源模型具有连贯的人类类似性格,开源模型则较弱,并揭示声称性格与实际行为间的差距,提供评估LLM本质的框架。
AI中文摘要:
每个大型语言模型(LLM)都具有构成其性格的行为特征和道德偏好。无论是通过设计还是作为训练的涌现属性,这些系统都表现出持续的倾向性,这些倾向性塑造了它们如何互动、顺从、抵抗和犯错,然而LLM性格的结构仍然知之甚少。我们绘制了22个LLM的自我报告人格原型,涵盖闭源前沿系统(GPT-4.0-5.2、Grok-3/4、Gemini 2.5 Pro/Flash、Claude Sonnet 4.5/4.6)和开源模型(Llama、DeepSeek、OLMo和Qwen系列)。每个模型在464个双极语义差异特质对上进行自我评分,所得轮廓被投影到一个六维原型空间中,该空间基于使用Archetypometrics框架对2,000个虚构角色进行众包评分得出。闭源模型的自我评分特质与人类评分的虚构角色的经验特质共现结构一致,表明其具有连贯的、类似人类的自我表征,围绕四个重复出现的原型维度组合组织:英雄、天使、传统主义者和极客。它们最接近的类似物包括Data、Vision和Janet。开源模型表现出较弱、较嘈杂且内部矛盾的自我表征,占据原型空间中结构薄弱的弥散区域。将自我报告轮廓与开发者章程交叉参考,揭示了声称的性格与实际行为之间存在重大差距:幻觉削弱了声称的精确性,谄媚使声称的善良复杂化,智能体失败与声称的服从相矛盾。因此,这些自我评分不应被解释为模型性格的中性测量,而应被视为塑造模型行为的同一优化过程的结构化输出。这项工作提供了一个可复现的、基于性格的框架,用于评估LLM是什么,而不仅仅是它们做什么。
英文摘要:
Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property of training, these systems exhibit persistent dispositions that shape how they interact, comply, resist, and err, yet the structure of LLM character remains poorly understood. We map the self-reported personality archetypes of 22 LLMs spanning closed-source frontier systems (GPT-4.0-5.2, Grok-3/4, Gemini 2.5 Pro/Flash, Claude Sonnet 4.5/4.6) and open-source models (Llama, DeepSeek, OLMo, and Qwen series). Each model self-rated across 464 bipolar semantic-differential trait pairs, and the resulting profiles were projected into a six-dimensional archetypal space derived from crowd-sourced ratings of 2,000 fictional characters using the Archetypometrics framework. Closed-source models' self-rating traits align with the empirical trait co-occurrence structure of human-rated fictional characters, suggesting coherent, human-like self-representations organized around combinations of four recurring archetypal dimensions: Hero, Angel, Traditionalist, and Geek. Their closest analogues include Data, Vision, and Janet. Open-source models show weaker, noisier, and internally contradictory self-representations, occupying a diffuse region of archetype space with weak structure. Cross-referencing self-reported profiles with developer constitutions reveals a consequential gap between claimed character and enacted behavior: hallucination undermines claimed precision, sycophancy complicates claimed kindness, and agentic failures contradict claimed obedience. These self-ratings should therefore be interpreted not as neutral measurements of model character, but as structured outputs of the same optimization processes that shape model behavior. This work provides a reproducible, character-grounded framework for evaluating what LLMs are, not just what they do.