发表机构
ByteDance Seed; Zhejiang University; National University of Singapore(字节跳动 Seed; 浙江大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出GST-Bench基准评估VLMs的视频全局空间感知,发现其与人类存在显著差距,构建GST-Bench-Local探究原因,还提供GST-Train数据集助力相关研究。
AI 中文摘要
空间智能是具身智能体的基础,但现有基准聚焦于从单一或少数视点的局部空间感知,忽略了对连续长时序视觉流的全局空间感知。为解决这一局限,我们提出全局时空基准(Global-Spatial-Temporal Benchmark,GST-Bench),这是一个用于视频理解中全局空间智能的视觉问答(VQA)基准,包含从6790分钟合成视频中生成的经人工验证的问题。该基准要求模型对输入视频中未见过的新视点进行准确的空间推理,并将自我中心视角的观测结果映射到全局顶视图图像上。我们对22个最先进的视觉语言模型(VLMs)进行的全面评估揭示了模型与人类之间存在显著差距:最强的零样本模型仅取得42.68分,远低于人类的79.08分。为探究该差距的原因,我们构建了GST-Bench-Local,发现尽管模型在相同任务设定下具备较强的局部空间理解能力,但仍无法将长时序观测结果整合为全局一致的场景表示。我们还提供了用于全局空间推理的数据集GST-Train,作为补充资源以推动未来针对该挑战的研究。
英文摘要
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce the Global-Spatial-Temporal Benchmark (GST-Bench), a VQA benchmark for global spatial intelligence in video understanding, comprising human-verified questions derived from 6,790 minutes of synthetically generated video. It requires models to perform accurate spatial inference from novel viewpoints unseen in the input video and to map egocentric observations onto global top-down images. A comprehensive evaluation of 22 state-of-the-art VLMs exposes a striking gap between models and humans: the strongest zero-shot model attains only 42.68, far below the human score of 79.08. To probe the cause of this gap, we construct GST-Bench-Local and find that models, despite strong local spatial understanding under the same task formulation, still fail to consolidate long-horizon observations into a globally consistent scene representation. We further provide GST-Train, a dataset for global spatial reasoning, as a complementary resource to facilitate future research on this challenge.