发表机构
The University of Hong Kong; National University of Singapore; Peking University(香港大学; 新加坡国立大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TempCloze视频时序推理基准,针对Video-LLMs的时序对齐瓶颈展开评估,通过多维度干扰项设计减少语言捷径,分析了模型错误的影响因素。
AI 中文摘要
针对视频-大语言模型(Video-LLMs)的时序推理基准常以语言为中介,存在选项措辞、答案关联或语言先验带来的语言捷径空间。为减少此类捷径,我们提出TempCloze——一款用于评估Video-LLMs视觉时序推理的视频完形填空基准。给定视频的起始和结束片段,模型必须从四个候选中识别出真实的缺失中间片段。TempCloze包含来自七个来源的1521个经精心筛选的视频,主要为长镜头和自我中心视频。我们从三个维度构建同源干扰项:语义维度询问应发生何种事件,对齐维度探测应在何时发生,进展维度测试应如何展开,而共享场景与对象则减少外观线索。我们对10个专有Video-LLMs和21个开源Video-LLMs的评估显示,对齐是主要瓶颈:模型常能识别合理的语义内容与局部事件进展,但在时序对齐上表现吃力。我们还对TempCloze-Mixed和TempCloze-Hard这两个子集,采用四个代表性模型开展错误模式与行为敏感性分析,以探究错误产生的位置,以及候选顺序、上下文方向、可见跨度、帧密度和测试时缩放如何影响模型选择。
英文摘要
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.
CommentsEMNLP 2026 Findings