发表机构
Tsinghua University; Shanghai Qi Zhi Institute; Carnegie Mellon University; Columbia University; Fangcun AI; University of Chinese Academy of Sciences; ShanghaiTech University(清华大学; 上海期智研究院; 卡内基梅隆大学; 哥伦比亚大学; 方寸智能; 中国科学院大学; 上海科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出AwarenessBench基准,从四个维度评估语言模型认知能力,发现先进模型表现优于随机基线,但多数在元认知和自我意识上落后于人类,且意识能力独立于语言建模进步。
AI 中文摘要
随着语言模型(LMs)展现出越来越多的类似意识的行为,评估其认知能力变得至关重要。我们引入了AwarenessBench,这是首个全面评估语言模型在四个维度认知能力的基准:元认知、自我意识、社会意识和情境意识,涵盖15种认知功能和14,381个样本。通过对18个最先进的语言模型进行评估,我们发现所有模型都持续超越随机基线,且更先进的模型表现更好。我们进一步将语言模型与三个不同人口统计学群体的人类表现进行比较,其中表现最好的模型在总体上超越了人类平均水平,但大多数模型在元认知和自我意识方面仍明显不足。最后,我们表明意识是一种独特的能力:语言建模或推理的进步并不必然转化为认知能力的提升。
英文摘要
As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social awareness, and situational awareness, covering 15 cognitive functions and 14,381 samples. Evaluating 18 state-of-the-art LMs, we find that all consistently surpass random baselines, with more advanced models performing better. We further compare LMs with human performance across three demographic groups, where the best-performing model surpasses human averages overall, but most still fall markedly short in metacognition and self-awareness. Finally, we show that awareness is a distinct capability: progress in language modeling or reasoning does not necessarily translate into improved cognition.