发表机构
School of Artificial Intelligence, Beihang University; Zhongguancun Academy; School of Engineering, Westlake University; School of Transportation Science and Engineering, Beihang University(北京航空航天大学人工智能学院; 中关村科学城; 西湖大学工学院; 北京航空航天大学交通科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究聚焦多模态大语言模型在动态场景中从连续视觉线索推理的能力,提出ViSTR-Bench评估套件,基于相关原则建立四维评估,含15个子任务和1340个问答对,经评估发现当前模型在复杂时空推理上有瓶颈,远不及人类。
AI 中文摘要
多模态大语言模型(MLLMs)在各种专家级任务中取得了显著成功,但在空间感知和动态推理等人类通过对现实世界的持续观察自然发展出的基本能力方面仍存在困难。近期研究认识到这一差距并引入专用基准来评估MLLMs的时空能力。然而,现有基准大多关注静态场景或需要精确的定量预测,对基于时间线索的直观推理探索不足。本文介绍了视觉空间 - 时间推理基准(ViSTR-Bench),这是一个旨在系统评估MLLMs能否从动态场景中的连续视觉线索进行定性推理的新颖评估套件。它基于时间强调、推理方向和定性评估原则,建立了涵盖运动感知、空间关系、结果预测和物理动力学的全面四维评估。该基准包含15个不同子任务和1340个高质量视频问答对,涵盖各种桌面、室内和室外场景。对多种最先进的专有、开源和专门的空间MLLMs的广泛评估表明,尽管当前模型具有强大的一般视频理解能力,但在复杂的时空推理中仍面临重大瓶颈,远低于人类表现。
英文摘要
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation of the real world, such as spatial perception and dynamic reasoning. Recent studies have recognized this gap and introduced dedicated benchmarks to evaluate the spatial-temporal capabilities of MLLMs. However, existing benchmarks mostly focus on static scenes or require exact quantitative predictions, leaving intuitive reasoning from temporal cues largely underexplored. In this paper, we introduce the Visual Spatial-Temporal Reasoning Benchmark (ViSTR-Bench), a novel evaluation suite designed to systematically assess whether MLLMs can perform qualitative reasoning from continuous visual cues in dynamic scenes. Guided by the principles of temporal emphasis, reasoning orientation, and qualitative evaluation, ViSTR-Bench establishes a comprehensive four-dimensional evaluations covering Motion Perception, Spatial Relations, Outcome Prediction, and Physical Dynamics. The benchmark comprises 15 distinct subtasks and 1,340 high-quality video question-answer pairs spanning diverse tabletop, indoor, and outdoor scenarios. Extensive evaluations of a broad spectrum of state-of-the-art proprietary, open-source, and specialized spatial MLLMs reveal that, despite their strong general video understanding capabilities, current models still face substantial bottlenecks in complex spatial-temporal reasoning and remain far below human performance.
Comments37 pages, 37 figures