arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26291cs.AIq-bio.NC

评估人类与大型语言模型的心智化能力

Assessing mentalization in humans and large language models

Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究以两项经济博弈结合认知计算建模,对比测试DeepSeek等四款LLM智能体与人类参与者,发现不同LLM心智化能力有差异,GPT-5表现优于人类,证实认知计算建模可用于评估人机比较智能。

中文摘要 AI 辅助

心智化(即推断他人信念与意图以指导自身选择的能力)是支撑人类社会互动的关键认知功能。大型语言模型(LLMs)在心理理论任务上展现出与人类一致的行为,然而这些模型是否能通过心智化引导适应性行为尚不明确。本研究采用两项结合认知计算建模的经济博弈,以揭示LLMs中心智化的潜在策略。我们测试了四个模型家族(DeepSeek、GPT-4.1、GPT-5、Gemini 2.0 Flash)的单个LLM智能体(样本量N=2099),对手为不同复杂程度的智能体,同时检验旨在引出策略推理的提示策略是否能提升表现。我们以人类参与者(样本量N=251)的结果作为对比基准。在两项博弈中,LLMs展现出清晰的心智化行为与计算特征,且这些特征因模型提供商和规模存在显著差异。策略提示通常通过诱导更复杂的推理提升表现,但两项任务的受益程度有所不同。最后,GPT-5智能体能灵活调整自身的递归推理深度以应对日益复杂的对手,表现优于人类参与者。综上,我们证明不同LLMs具备不同的心智化能力,并强调认知计算建模是评估人类与机器比较智能的正式方法。

英文摘要

Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.

发表机构

  • University of Birmingham(伯明翰大学)
  • Virginia Tech(弗吉尼亚理工大学)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑