arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39238cs.CL

4MT-VLM:视觉语言模型的认知地图有多粗糙?

4MT-VLM: How Coarse Is a VLMs Cognitive Map?

Markus Frey

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出4MT-VLM数据集,测试16种视觉语言模型在视角变化下的地点识别能力,发现其空间认知地图粗糙,旋转135°时性能低于随机水平,远逊于人类。

中文摘要 AI 辅助

一个移动的智能体必须从从未见过的视角识别一个地点。我们引入了4MT-VLM,一个程序化生成景观的数据集,每个景观以五种刺激模式渲染,在保持布局固定的同时去除外观线索:形状和颜色、仅形状、仅颜色、无物体的裸地形山峰,以及将山峰置于地平线上的山谷视角。最后一种条件在临床上常用于探测人类患者的 hippocampal 功能。我们在十六个不同的开源和闭源模型上测试该基准,并报告4AFC性能,这一指标也用于对人类参与者进行评分。我们观察到,模型能从研究过的视角识别地点,但一旦相机移动,它们便失去该能力,在135°处降至25%的随机水平以下,而人类观察者在此得分85%。前沿模型(Gemini 3.8 Flash、GPT-5.6)仅正确回答39%和31%的旋转试验,仅当干扰物相距超过30米时才恢复到85%和55%。我们的基准表明,尽管当前视觉语言模型具备初步的认知地图,但其空间分辨率从根本上仍过于粗糙,无法在视角变化时维持稳定的3D世界理解。

英文摘要

An agent that moves must recognise a place from a viewpoint it has never seen. We introduce 4MT-VLM, a dataset of procedurally generated landscapes, each rendered across five stimulus modes that remove appearance cues while holding layout fixed: shape and colour, shape only, colour only, bare terrain peaks with no objects, and a valley viewpoint that puts the peaks on the horizon. The last condition is commonly used in clinics to probe hippocampal function in human patients. We test this benchmark across sixteen different open and closed-source models and report 4AFC performance, a measure which is also used to grade human participants. We observe that models identify a place from the studied viewpoint but lose it once the camera moves, dropping below the 25% chance level at 135° where a human observer scores 85%. Frontier models (Gemini 3.8 Flash, GPT-5.6) answer only 39% and 31% of rotated trials correctly, recovering to 85% and 55% only when distractors are moved more than 30 meters apart. Our benchmark demonstrates that while current VLMs possess rudimentary cognitive maps, their spatial resolution remains fundamentally too coarse to maintain a stable, 3D understanding of the world once the viewpoint changes.

发表机构

  • Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习和人工智能研究所)
  • Fraunhofer IAIS, University of Bonn(弗劳恩霍夫智能分析和信息系统研究所,波恩大学)

机构由 AI 辅助整理,请以论文原文为准。

↑