arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KilometerVision:视觉语言模型中大规模空间智能的新前沿

KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs

Aravindh Mahendran, Michael King, Matthew Koichi Grimes, Antoine Yang, Tyler Zhu, Joseph Heyward, Tengda Han, Shiry Ginosar, Chen Sun, Dima Damen, Simon Osindero, Noah Snavely, Simon Lynen, João Carreira, Viorica Pătrăucean

arXiv 2609.39588首次发表:更新:

发表机构

Google DeepMind; Princeton University; Toyota Technological Institute at Chicago(谷歌DeepMind; 普林斯顿大学; 芝加哥丰田技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出KilometerVision基准,首次从真实视频探测VLM在1公里尺度上的地理布局理解,发现模型依赖2D识别和文本匹配而非真正空间推理。

AI 中文摘要

我们推动了视觉语言模型(VLMs)中大规模空间智能的前沿,并引入了首个从真实世界视频中探测地理布局理解的基准,其跨度可达1公里。受认知科学文献启发,我们根据人类空间意识的层次阶段来评估模型:通过地标进行锚定,通过路线连接它们,并将这些整合为全局心理地图。大量实验揭示了当前AI模型处理空间信息时的根本性分歧。我们发现,视觉语言模型并非利用真正的路径整合或形成几何测量知识,而是几乎完全依赖二维视觉识别和文本匹配来绕过复杂的空间推理。该基准可在此https URL公开获取。

英文摘要

We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark is publicly available at https://perception-test-challenge.github.io/kilometervision.html.

Journal refECCV 2026, 2026, pages 194--212

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑