arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00960cs.CV

Video-Index:一个精选的视频理解元基准

Video-Index: A Curated Meta-Benchmark for Video Understanding

Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu

首次发表
浏览论文内容

中文总结 AI 辅助

我们提出Video-Index,一个基于攻击金字塔审计115个基准并筛选出840个最难验证条目的视频理解元基准,揭示现有基准的捷径漏洞,并展示Claude Opus 5显著领先开源模型。

中文摘要 AI 辅助

一个视频基准应该奖励它所声称衡量的能力,然而模型可以利用答案选项、问题文本或部分视觉证据。我们引入了攻击金字塔,即五个级别的捷径攻击,每个级别对每个条目的访问权限逐渐增加,并使用它审计了115个视频基准。在35个基准上,从未见过帧的攻击者接近全视频准确率。在51个具有时间探测的基准上,打乱帧保持了全视频准确率的中位数96%。在63个基准中,近似重复的问题至少占条目的一半。我们从其中112个基准中筛选了505,518个问答对,进入一个经过审计的池子。智能体将评估请求转化为规范,一个带有红队门的确定性选择器组合出可复现的基准。我们发布了Video-Index,即在四个能力组中每个组内经过这些攻击的最难的210个已验证条目,共840个条目,来自76个来源。在相同的固定输入下,Claude Opus 5的得分超过每个开源模型超过37个百分点,智能体工具额外增加约20个百分点,但所有系统在效率和准确性上仍有提升空间。博客:此HTTPS URL GitHub:此HTTPS URL Hugging Face:此HTTPS URL

英文摘要

A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers that never see a frame approach full-video accuracy. On 51 benchmarks with temporal probes, shuffled frames keep a median 96% of full-video accuracy. Near-duplicate questions make up at least half the items in 63 benchmarks. We screen 505,518 question-answer pairs from 112 of them into an audited pool. Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks. We release Video-Index, the 210 hardest verified items under these attacks in each of four capability groups, 840 items from 76 sources. With the same fixed input, Claude Opus 5 outscores every open-source model by over 37 percentage points, and agent tools add about 20 more, yet all systems leave room to improve efficiency and accuracy. Blog: https://www.enxinsong.com/blog/video-index/ GitHub: https://github.com/Espere-1119-Song/Video-Index Hugging Face: https://huggingface.co/datasets/Video-Index/Video-Index

发表机构

  • University of Pennsylvania(宾夕法尼亚大学)
  • New York University(纽约大学)
  • Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑