arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低频陷阱:视频语言模型在简单事件记录上的失败

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

Sarvesh Baskar, Zikui Cai, Shayan Shabihi, Anirudh Satheesh, Muhammad R. Islam, Udari Madhushani Sehwag, Tom Goldstein, Furong Huang

arXiv 2608.06361首次发表:更新:

AI 中文总结

该研究提出基于轨迹的参数化分析方法,发现视频语言模型Gemini 3.6 Flash在高数量、高频率事件计数上表现极差,额外帧和提示策略难以解决问题,推动视频评估转向时间推理故障诊断。

AI 中文摘要

现实世界的视频基准测试覆盖范围广泛,但它们的固定片段会将事件数量、速率、持续时间和视觉复杂性交织在一起,使得故障模式难以分离。虽然现有的程序化基准测试提供了更好的控制,但它们仅对最终答案进行评分,而不是将报告的事件与可执行的真实情况进行审计。为了弥合这一差距,我们引入了基于轨迹的参数化分析,用于三个受控视频任务中的事件计数:弹跳球的壁接触、视觉眨眼和分类状态转换。在2190个视频中,我们在保持渲染固定的同时改变事件数量N和频率F。每个视频都包含一个可执行的事件轨迹,用于能力表面估计和时间戳级别的评估。我们的结果揭示了阶段性的时间故障:在80%的可靠性阈值下,Gemini 3.6 Flash能够可靠地计数0.5和1.0 Hz下最多12个的持续状态转换,但对于瞬态眨眼事件却没有可靠的正计数区域。因此,事件表示决定了模型是否最初访问证据——这一限制会随着数量和频率的增加而加剧。在高数量、高频率的 regime 中,最终计数的正确率仅为0.2%,模型仅恢复了18.1%的真实事件。为了测试视觉访问是否是主要瓶颈,我们提高了采样率。尽管这将弹跳球的准确率从19.6%提高到了29.3%,但报告的序列与真实情况一致的时间仅为3.7%。因此,额外的帧可以提高最终得分,而不会产生忠实的事件恢复。不同的提示策略也产生了类似的有限增益,现实世界的视频评估显示成功同样集中在低事件数量上。最终,基于轨迹的分析将视频评估从聚合准确率指标转变为对时间推理失败位置的详细诊断。

英文摘要

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled video tasks: bouncing-ball wall contacts, visual blinks, and categorical state transitions. Across 2,190 videos, we vary event count N and frequency F while holding rendering fixed. Each video includes an executable event trace for capability-surface estimation and timestamp-level evaluation. Our results reveal a staged temporal failure. At an 80% reliability threshold, Gemini 3.6 Flash reliably counts persistent state transitions up to 12 events at 0.5 and 1.0 Hz, yet demonstrates no reliable positive-count region for transient blinking events. Thus, event representation dictates whether a model initially accesses evidence -- a limitation that compounds as count and frequency increase. In the high-count, high-frequency regime, only 0.2% of final counts are correct and the model recovers just 18.1% of true events. To test if visual access is the primary bottleneck, we increase sampling rate. Although this boosts Bounce Ball accuracy from 19.6% to 29.3%, the reported sequence agrees with ground truth only 3.7% of the time. Extra frames can therefore inflate final scores without producing faithful event recovery. Different prompting strategies yield similarly limited gains, and real-world video evaluations show the same concentration of success at low event counts. Ultimately, trace-grounded profiling shifts video evaluation from aggregate accuracy metrics to a detailed diagnostic of where temporal reasoning fails.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑