arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.00288cs.CVcs.AI

TimeBlind: 一种用于视频大语言模型的时空组合性基准

TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs

  • University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
  • University of Pittsburgh(匹兹堡大学)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra, Jean de Dieu Nyandwi, Gedas Bertasius

AI总结:

TimeBlind通过细粒度时空理解基准揭示视频大语言模型对时间动态理解的不足,展示其在实例准确率上的显著劣势,强调对静态视觉捷径的依赖。

AI中文摘要:

细粒度的时空理解对于视频推理和具身AI至关重要。然而,尽管多模态大语言模型(MLLMs)擅长静态语义,但它们对时间动态的理解仍然脆弱。我们提出了TimeBlind,一种用于组合性时空理解的诊断基准。受认知科学的启发,TimeBlind将细粒度时间理解分为三个层次:识别原子事件、描述事件属性以及推理解事件之间的依赖关系。与将识别与时间推理混为一谈的基准不同,TimeBlind利用最小对法范式:视频对共享相同的静态视觉内容,但仅在时间结构上有所不同,利用互补问题来中和语言先验。在20种最先进的MLLMs(如GPT-5、Gemini 3 Pro)上评估了600个精心挑选的实例(2400个视频-问题对),结果显示,表现最好的MLLM的实例准确率(正确区分一对视频)仅为48.2%,远低于人类表现(98.2%)。这些结果表明,即使前沿模型也严重依赖静态视觉捷径而不是真正的时序逻辑,将TimeBlind定位为下一代视频理解的重要诊断工具。数据集和代码可在https://baiqi-li.github.io/timeblind_project/上获得。

英文摘要:

Fine-grained spatio-temporal understanding is essential for video reasoning and embodied AI. Yet, while Multimodal Large Language Models (MLLMs) master static semantics, their grasp of temporal dynamics remains brittle. We present TimeBlind, a diagnostic benchmark for compositional spatio-temporal understanding. Inspired by cognitive science, TimeBlind categorizes fine-grained temporal understanding into three levels: recognizing atomic events, characterizing event properties, and reasoning about event interdependencies. Unlike benchmarks that conflate recognition with temporal reasoning, TimeBlind leverages a minimal-pairs paradigm: video pairs share identical static visual content but differ solely in temporal structure, utilizing complementary questions to neutralize language priors. Evaluating over 20 state-of-the-art MLLMs (e.g., GPT-5, Gemini 3 Pro) on 600 curated instances (2400 video-question pairs), reveals that the Instance Accuracy (correctly distinguishing both videos in a pair) of the best performing MLLM is only 48.2%, far below the human performance (98.2%). These results demonstrate that even frontier models rely heavily on static visual shortcuts rather than genuine temporal logic, positioning TimeBlind as a vital diagnostic tool for next-generation video understanding. Dataset and code are available at https://baiqi-li.github.io/timeblind_project/ .

补充信息

↑