有何玄机?评估视觉-语言模型的时间一致性
What's the Catch? Evaluating Temporal Consistency in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出TimeCatch基准,发现视觉-语言模型可检测单帧异常但难整合跨帧信息,在时间异常检测上表现接近随机水平。
中文摘要 AI 辅助
视觉-语言模型(VLMs)在视频和图像序列基准测试中表现出色,但目前尚不清楚它们是否能捕捉时间结构。为研究这一问题,我们将时间定位问题建模为异常检测问题,提供了一种简单且可控的评估方法,可直接测试模型对时间一致性的敏感性。我们提出了TimeCatch方法,通过交换连续帧创建时间异常,以及用高斯噪声替换单帧创建帧级异常。在四项合成和真实数据集上,针对异常检测和定位任务对模型进行评估,同时开展了人类对照实验。我们的评估显示,帧级异常检测与时间异常检测之间存在显著差距:VLMs能一致检测到帧级异常,且通常能准确定位;但在时间异常检测上表现接近随机水平,定位仅略高于随机水平。相比之下,人类在两项任务上均达到接近天花板的性能。对模型规模、提示策略、序列长度和视觉相似度的进一步分析表明,这些失败无法仅用感知或模型容量的限制来解释。综合来看,这些发现表明,当前VLMs能够识别单帧内的异常,但难以整合跨帧信息以推理时间一致性。TimeCatch为评估视觉-语言模型的时间定位提供了可控基准。
英文摘要
Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection problem, providing a simple and controlled evaluation that directly tests sensitivity to temporal consistency. We introduce TimeCatch, where temporal anomalies are created by swapping consecutive frames and frame-level anomalies by replacing a frame with Gaussian noise. Models are evaluated on anomaly detection and localization tasks across four synthetic and real-world datasets, alongside a human study. Our evaluation reveals a substantial gap between frame-level and temporal anomaly detection. While VLMs consistently detect frame-level anomalies and often localize them accurately, under our main evaluation setting they generally perform near chance on temporal anomaly detection and show limited localization performance. Humans, in contrast, achieve near-ceiling performance on both tasks. Additional analyses across model scales, prompting strategies, sequence lengths, and visual similarity show that performance can improve under some conditions, while substantial gaps in temporal anomaly detection and localization remain. Together, these findings reveal a gap between frame-level and temporal anomaly detection. TimeCatch provides a controlled benchmark for evaluating temporal consistency in vision-language models.
发表机构
- University of Copenhagen(哥本哈根大学)
机构由 AI 辅助整理,请以论文原文为准。