发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究评估9个前沿VLM在两项心理学衍生ToM任务中的表现,发现其ToM特征存在跨任务解离,无模型在两项任务中均接近典型发展成人,推理可改善部分模型在导演任务中的表现。
AI 中文摘要
前沿视觉语言模型(VLM)在不同任务中是否呈现与同一人类参考组匹配的连贯心智理论(ToM)特征,还是该特征会从一种范式碎片化到另一种?我们在两个心理学衍生基准上评估了一组共9个前沿VLM:基萨尔导演任务(自我中心干扰下的视觉观点采择)和弗里思-哈佩动画三角形任务(用卡斯特利评分标准从纯运动中归因意图)。在导演任务中,无思维链时,该组模型在78%的试次中出现自我中心错误,类似儿童而非成人;模型间差异显著,推理可改善部分模型表现。在三角形任务中,该组模型意图归因不足:其ToM特征与高功能自闭症成人(HF-ASD)均值的距离是典型发展成人(TD)均值的三倍多,而目标导向模型和随机模型则接近TD。没有模型在两项任务中都最接近TD:在导演任务中表现得像成人的模型,在三角形任务中处于HF-ASD一侧;在三角形任务中最像TD的模型,在导演任务中则像儿童。我们报告的是群体层面描述,而非任何模型的诊断标签。
英文摘要
Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nine frontier VLMs on two psychology-derived benchmarks: the Keysar Director Task (visual perspective-taking under egocentric interference) and the Frith-Happé animated triangles scored with the Castelli rubric (intention attribution from pure motion). On the Director Task, without chain-of-thought, the panel makes the egocentric error on 78\% of trials like children rather than adults; variation is substantial across models, and reasoning rescues several models. On the triangles, the panel under-attributes intention: its ToM profile sits more than three times closer to the high-functioning-autistic-adult (HF-ASD) mean than to the typical-development-adult (TD) mean, while Goal-Directed and Random stay near TD. No model is nearest TD on both tasks; the model that looks adult-like on the Director Task falls on the HF-ASD side on the triangles, and the most TD-like model on the triangles is child-like on the Director Task. We report group-level descriptions, not diagnostic labels for any model.