arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28991cs.CVcs.AI

分数之下:重新思考视频理解模型的幻觉评估

Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models

Shuzhi Gong, Fengze Sun, Yuansan Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出因果阶段干预协议,通过60,008次运行揭示视频理解多阶段LLM智能体中定位是幻觉的主要来源,其影响约为视觉观察的四倍,并指出错误证据比缺失证据危害更大,现有基准分数无法可靠预测因果级联敏感性。

中文摘要 AI 辅助

视频理解越来越多地由多阶段LLM智能体执行,这些智能体将时间定位、视觉观察和推理分离。然而,这些阶段通常在不同的基准和分布上进行评估,使得难以确定幻觉源于何处。我们首先围绕这些阶段组织现有基准,并表明它们的分数提供不一致的诊断信号:更强的阶段级性能并不可靠地意味着更低的下游幻觉,甚至针对相同能力的基准也可能存在分歧。因此,我们引入了一种因果阶段干预协议,该协议在保持下游任务固定的同时覆盖各个阶段。在三种视频智能体架构的60,008次运行中,我们发现定位是下游错误的主要来源,其因果影响大约是破坏视觉观察的四倍。成功的定位主要依赖于定位正确的区域,而非精确的时间重叠,这解释了为什么标准的mIoU指标难以预测下游可靠性。我们进一步发现,错误的证据比缺失的证据危害大得多。最后,针对这些干预措施审计现有基准表明,它们的分数不能可靠地预测因果级联敏感性,并且可能在分布偏移下失效。这些结果激励了基于干预的、阶段感知的评估,以构建可信赖的视频智能体。

英文摘要

Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing benchmarks around these stages and show that their scores provide inconsistent diagnostic signals: stronger stage-level performance does not reliably imply lower downstream hallucination, and even benchmarks targeting the same capability can disagree. We therefore introduce a causal stage-intervention protocol that overwrites individual stages while holding the downstream task fixed. Across 60,008 runs on three video-agent architectures, we find that grounding is the dominant source of downstream error, with roughly four times the causal impact of corrupting visual observations. Successful grounding depends primarily on locating the correct region rather than precise temporal overlap, explaining why standard mIoU metrics poorly predict downstream reliability. We further find that incorrect evidence is substantially more harmful than missing evidence. Finally, auditing existing benchmarks against these interventions reveals that their scores do not reliably predict causal cascade sensitivity and can fail under distribution shift. These results motivate intervention-based, stage-aware evaluation for trustworthy video agents.

发表机构

  • The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑