发表机构
HKUST (GZ); Beta Infinity; Beihang University; UESTC; Tsinghua University(香港科技大学(广州); Beta Infinity; 北京航空航天大学; 电子科技大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
STORM-Bench提出在线视频问答基准,含5736个问题,通过STORM-BR指标揭示视频大语言模型在不确定查询上的过度自信,暴露准确率掩盖的认知可靠性缺陷。
AI 中文摘要
可靠的在线视频问答需要在跟踪状态转换的同时,在视觉证据不足时选择性地弃权(不执行)。现有基准侧重于静态识别或长程检索,很少在演变和不完整证据下评估这些耦合能力。我们提出了STORM-Bench,包含5,736个问题,覆盖630个紧凑、变化密集的情节,跨越五个以自我为中心的领域(STORM-Real)和两个受控模拟子集(STORM-Sim),帧率为1 FPS。问题根据累积变化强度的代理(低、中、高)和查询时的可回答性(已知、不确定)进行分层。为了衡量可靠性,我们引入了STORM-BR,这是一个基于联合答案状态正确性的调和指标,能够揭示被总体准确率掩盖的弃权失败,以及用于不确定性归因的STORM-BR-ATTR。在14个视频大语言模型中,在线准确率峰值达到60.3%(平均51.7%),而STORM-BR的范围从5.7%到35.6%(平均18.8%),这主要是由于对不确定查询的普遍过度自信。STORM-Bench表明,任务准确率掩盖了这些认知可靠性和状态跟踪方面的差距。基准和代码可在以下网址获取:此https URL。
英文摘要
Reliable online video question answering requires tracking state transitions while selectively abstaining when visual evidence is insufficient. Existing benchmarks focus on static recognition or long-range retrieval, rarely evaluating these coupled capabilities under evolving and incomplete evidence. We present STORM-Bench, comprising 5,736 questions across 630 compact, change-dense episodes spanning five egocentric domains (STORM-Real) and two controlled simulation subsets (STORM-Sim) at 1 FPS. Questions are stratified by a proxy for accumulated change intensity (Low, Medium, High) and query-time answerability (Known, Uncertain). To measure reliability, we introduce STORM-BR, a harmonic metric over joint answer-status correctness that exposes abstention failures masked by aggregate accuracy, alongside STORM-BR-ATTR for uncertainty attribution. Across 14 video LLMs, online accuracy peaks at 60.3\% (mean 51.7\%), whereas STORM-BR ranges from 5.7\% to 35.6\% (mean 18.8\%), driven by pervasive overconfidence on uncertain queries. STORM-Bench shows that task accuracy masks these gaps in epistemic reliability and state tracking. Benchmark and code are available at https://github.com/siruzhong/STORM-Bench.
Comments50 pages, 19 figures, 27 tables