超越莎莉-安妮任务:使用认知谢林点评估大语言模型中的心理理论
Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points
- Paradigms of Intelligence Team(智能范式团队)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究旨在解决大语言模型心理理论评估问题,引入认知不对称谢林任务,通过要求模型在不同认知透明度下收敛于语义谢林点来评估其功能性ToM能力,发现能力差距及协调失败原因,指出社会推理和认知跟踪是瓶颈,为LLM评估与发展提供目标。
AI中文摘要:
基于文本对大语言模型(LLMs)中心理理论(ToM)的评估,常涉及类似莎莉-安妮任务的认知测试,因预训练中接触相似任务可能被破解,且不能有效测试模型在自然场景下的功能性ToM能力。为解决这些问题,我们引入认知不对称谢林任务(EAST),这是一个两人对话游戏,用于评估强大且可推广的ToM能力。通过要求LLM-LLM二元组在不同认知透明度状态下独立收敛于语义谢林点,评估模型是否能稳健应用ToM实现协调。结果揭示了功能性社会推理方面的显著能力差距,只有前沿模型能成功应对任务中不同的认知需求。推理轨迹分析表明,协调失败主要由认知跟踪错误导致。尽管在传统静态基准测试中表现出色,但强大的社会推理和认知跟踪仍是关键瓶颈,为未来LLM评估和发展提供了具体目标。
英文摘要:
Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models' functional ToM abilities in ways that generalize to naturalistic settings. To address these issues, we introduce the Epistemic Asymmetry Schelling Task (EAST), a two-player dialogue game designed to benchmark robust and generalizable ToM abilities. By requiring LLM-LLM dyads to independently converge on semantic Schelling points under varying states of epistemic transparency, we evaluate whether models can robustly apply ToM to achieve coordination. Our results reveal a significant capability gap in functional social reasoning, with only frontier models successfully navigating the varying epistemic demands of the tasks. Analysis of reasoning traces shows that coordination failures are primarily driven by epistemic tracking errors, such as conflating private knowledge with mutual knowledge. Despite high performance on traditional static benchmarks, our study shows that robust social reasoning and epistemic tracking remain a critical bottleneck, providing concrete targets for future LLM evaluation and development.