CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models
CogToM:一种受人类认知启发的大型语言模型全面理论思维基准
机构 * BrainCog Lab, Institute of Automation, Chinese Academy of Sciences(脑认知实验室,自动化研究所,中国科学院) ; Beijing Institute of AI Safety and Governance (Beijing-AISI)(北京人工智能安全与治理研究所) ; Beijing Key Laboratory of Safe AI and Superalignment(北京安全人工智能与超对齐重点实验室) ; School of Artificial Intelligence, UCAS(人工智能学院,中国科学院大学) ; Long-term AI(长期人工智能)
专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI
AI总结 CogToM通过8000个双语实例评估LLM的理论思维能力,揭示性能异质性和认知瓶颈,为研究LLM认知边界提供新视角。