从视觉到预见:视觉语言模型中的预测性空间推理
From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
针对现有视觉语言模型缺乏预测性空间推理能力的问题,提出度量尺度模型SpatialMind,通过度量深度适配器和渐进状态链实现未来帧距离、运动方向及空间关系预测,并构建SpatialMind-30K数据集和SpatialMind-2K基准,实验显示其性能显著优于现有模型。
中文摘要 AI 辅助
预测未来空间状态有助于在动态环境中进行碰撞避免和及时决策。然而,现有的视觉语言模型(VLM)和空间推理基准主要关注已观测场景,对超出观测区间之外的预测性空间推理探索不足。为此,我们引入了SpatialMind,一个用于空间推理和未来预测的度量尺度视觉语言模型。其度量深度适配器将空间推理锚定到真实世界尺度,而其渐进状态链则建立当前空间状态和观测动态作为未来预测的基础。给定一个视频前缀,SpatialMind能够预测已观测和未观测的未来帧中的距离、运动方向和空间关系。为了训练和评估,我们构建了一个可扩展的数据引擎,该引擎将实体描述锚定在度量几何中,以生成问答对和状态监督。利用该引擎,我们构建了SpatialMind-30K数据集和SpatialMind-2K基准,两者均涵盖驾驶和日常第一人称视角场景。该基准涵盖三个层次的八项任务:当前状态理解、观测动态理解和未来预测。实验表明,SpatialMind在我们的基准上显著优于通用模型和空间专用模型,同时在VSI-Bench、OSI-Bench和VLM4D上取得了具有竞争力的零样本性能。
英文摘要
Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatial reasoning beyond the observed interval underexplored. To this end, we introduce SpatialMind, a metric-scale VLM for spatial reasoning and future prediction. Its metric depth adapter anchors spatial reasoning to real-world scale, while its progressive state chain establishes current spatial states and observed dynamics as the foundation for future prediction. Given a video prefix, SpatialMind predicts distances, motion directions, and spatial relations in both observed and unseen future frames. For training and evaluation, we build a scalable data engine that grounds entity descriptions in metric geometry to generate question-answer pairs and state supervision. Using this engine, we construct the SpatialMind-30K dataset and the SpatialMind-2K benchmark, both covering driving and everyday egocentric scenes. The benchmark spans eight tasks across three levels: current-state understanding, observed-dynamics understanding, and future prediction. Experiments show that SpatialMind substantially outperforms both general and spatially specialized models on our benchmark while achieving competitive zero-shot performance on VSI-Bench, OSI-Bench, and VLM4D.
发表机构
- University of Illinois at Chicago(伊利诺伊大学芝加哥分校)
- Bosch Research North America(博世北美研究院)
- Bosch Center for Artificial Intelligence (BCAI)(博世人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。