发表机构
National Taiwan University; NVIDIA(国立台湾大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ActionLens提出一个含6,701个视频问题的诊断基准,系统评估视觉语言模型在时空绑定上的失败模式,发现模型在动作与人物关联上存在系统性错误,并量化了数值解析与视觉框的差距。
AI 中文摘要
视频视觉语言模型在流行基准测试中得分超过80%,但在时空绑定方面仍存在困难:即在正确的时刻将正确的动作与正确的人关联起来。我们引入了ActionLens,一个包含6,701个多项选择视频问题的诊断基准,涵盖五个针对性诊断:转换检测、特定演员识别、并发动作绑定、定向交互推理和注视检测。真实答案由158万条每秒、每人的注释确定性推导而来。十四轮人工质量工程将答案清晰度从53%提升至超过90%的人工准确率。在20个视觉语言模型中,全集领先者得分为68.8%;在人工审核的子集上,其得分为65.9%,而汇总人工参考得分为91.0%。注视检测仍接近随机水平,而人工准确率为89.6%。在演员消歧方面,参考接口控制显示,关系描述比静态坐标恢复5.55至13.25个百分点,证实了显著的数值解析惩罚;然而,视觉框仍领先每个模型1.15至6.50个百分点,暴露出残余的无框演员解析差距。绑定陷阱分析显示,模型系统性地选择错误演员的动作。ActionLens为不同模型家族和规模提供了这些不同失败模式的诊断测量,以便直接比较。我们在此https URL发布所有数据、代码和评估脚本。
英文摘要
Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Ground-truth answers are derived deterministically from 1.58 million per-second, per-person annotations. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Across 20 VLMs, the full-set leader scores 68.8%; on the human-reviewed subset, it scores 65.9% versus 91.0% for the pooled human reference. Gaze detection remains near chance against 89.6% human accuracy. On actor disambiguation, reference-interface controls show that relational descriptions recover 5.55--13.25 points over static coordinates, confirming a substantial numeric-parsing penalty; yet visual boxes still lead every model by 1.15--6.50 points, exposing a residual unboxed actor-resolution gap. A binding-trap analysis shows models systematically select the wrong actor's action. ActionLens provides diagnostic measurements of these distinct failure modes across model families and scales for direct comparison. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276
CommentsProject Page: https://joslefaure.github.io/actionlens/