发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视频语言模型基准测试中视觉依赖问题,引入视觉依赖差距(VDG),通过实验表明准确性与视觉依赖可分离,任务类型排名稳定,帧多样性贡献大,H.264实验有新发现,VDG可作标准审查,推动相关研究。
AI 中文摘要
视频大语言模型(LLMs)中的基准准确性常被视为视觉理解的证据。我们对20个参数在2 - 78B之间、涵盖10种架构系列的模型进行了审查。引入视觉依赖差距(VDG),即原始视频与黑屏条件下每题正确性的差异。在MVBench上进行配对McNemar测试表明准确性和视觉依赖是可分离的。跨模型任务类型排名稳定,从黑屏到单帧、混洗帧和原始视频的诊断阶梯揭示了帧多样性提供了大部分视觉益处,时间顺序贡献近零准确性。H.264实验表明稳定的总体准确性掩盖了双向问题级答案翻转。该诊断也适用于四个通过API访问的模型。这些结果促使将VDG作为视频基准是否衡量视觉基础能力的标准审查。
英文摘要
Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty models spanning 2-78B parameters and ten architecture families. We introduce the Visual Dependency Gap (VDG), the difference in per-question correctness between original-video and black-screen conditions. Paired McNemar tests on MVBench show that accuracy and visual dependency are separable: models differ on original video (p = 0.0003) but not on black screens (p = 0.53). Across models, task-type rankings are stable: Attribute Perception is strongly visual, whereas Temporal Reasoning approaches the language-only baseline. A diagnostic ladder from black screen to single frame, shuffled frames, and original video reveals that frame diversity supplies most of the visual benefit, while temporal order contributes near-zero accuracy across sixteen open-weight models. An ablation from 0.5 to 24 FPS rules out sparse sampling as the cause. H.264 experiments further show that stable aggregate accuracy conceals bidirectional question-level answer flips. The diagnostic also generalizes to four API-accessed models, whose VDG values range from 0.025 to 0.315. These results motivate VDG as a standard audit for whether video benchmarks measure visually grounded capability. Code is available at https://github.com/JaeLee18/accuracy-without-grounding.
CommentsAccepted, ACM International Conference on Multimedia 2026 (ACM MM)