arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STORM-Bench:在演变与不完整证据下的在线视频问答评估

STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence

Siru Zhong, Shenghan Tan, Rihong Yan, Xiaohui Lv, Yuzheng Zhuang, Shuai Tao, Wulong Liu, Haohuan Fu, Yuxuan Liang

arXiv 2609.30981首次发表:更新:

发表机构

HKUST (GZ); Beta Infinity; Beihang University; UESTC; Tsinghua University(香港科技大学(广州); Beta Infinity; 北京航空航天大学; 电子科技大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

STORM-Bench提出在线视频问答基准,含5736个问题,通过STORM-BR指标揭示视频大语言模型在不确定查询上的过度自信,暴露准确率掩盖的认知可靠性缺陷。

AI 中文摘要

可靠的在线视频问答需要在跟踪状态转换的同时,在视觉证据不足时选择性地弃权(不执行)。现有基准侧重于静态识别或长程检索,很少在演变和不完整证据下评估这些耦合能力。我们提出了STORM-Bench,包含5,736个问题,覆盖630个紧凑、变化密集的情节,跨越五个以自我为中心的领域(STORM-Real)和两个受控模拟子集(STORM-Sim),帧率为1 FPS。问题根据累积变化强度的代理(低、中、高)和查询时的可回答性(已知、不确定)进行分层。为了衡量可靠性,我们引入了STORM-BR,这是一个基于联合答案状态正确性的调和指标,能够揭示被总体准确率掩盖的弃权失败,以及用于不确定性归因的STORM-BR-ATTR。在14个视频大语言模型中,在线准确率峰值达到60.3%(平均51.7%),而STORM-BR的范围从5.7%到35.6%(平均18.8%),这主要是由于对不确定查询的普遍过度自信。STORM-Bench表明,任务准确率掩盖了这些认知可靠性和状态跟踪方面的差距。基准和代码可在以下网址获取:此https URL。

英文摘要

Reliable online video question answering requires tracking state transitions while selectively abstaining when visual evidence is insufficient. Existing benchmarks focus on static recognition or long-range retrieval, rarely evaluating these coupled capabilities under evolving and incomplete evidence. We present STORM-Bench, comprising 5,736 questions across 630 compact, change-dense episodes spanning five egocentric domains (STORM-Real) and two controlled simulation subsets (STORM-Sim) at 1 FPS. Questions are stratified by a proxy for accumulated change intensity (Low, Medium, High) and query-time answerability (Known, Uncertain). To measure reliability, we introduce STORM-BR, a harmonic metric over joint answer-status correctness that exposes abstention failures masked by aggregate accuracy, alongside STORM-BR-ATTR for uncertainty attribution. Across 14 video LLMs, online accuracy peaks at 60.3\% (mean 51.7\%), whereas STORM-BR ranges from 5.7\% to 35.6\% (mean 18.8\%), driven by pervasive overconfidence on uncertain queries. STORM-Bench shows that task accuracy masks these gaps in epistemic reliability and state tracking. Benchmark and code are available at https://github.com/siruzhong/STORM-Bench.

Comments50 pages, 19 figures, 27 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑