arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21521cs.CVcs.AI

VidOmni-Bench:通过跨复杂度和时长的时空事件验证进行细粒度视频理解的基准

VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito, Donghyun Kim

首次发表
浏览论文内容

中文总结 AI 辅助

针对视频大语言模型细粒度理解评估难题,提出VidOmni-Bench基准,通过时空事件验证检测幻觉,发现模型在密集字幕和验证上均存在缺陷。

中文摘要 AI 辅助

尽管视频大语言模型(Video-LLMs)近期展现出强大性能,但可靠地评估其细粒度视频理解能力仍具挑战性。现有基准通常依赖问答或与真实字幕匹配,模型可能通过表面线索和不完整标注取得成功。为此,我们提出VidOmni-Bench,该基准要求模型验证密集视频字幕中的每个事件是否得到视频支持。VidOmni-Bench包含500个视频,涵盖五种复杂度类型,时长从4秒到90分钟不等。在沿这些维度收集视频后,我们使用多种Video-LLMs生成密集字幕,并获得人工验证的句子级标签,其中包含错误事件的句子作为评估的硬负样本。我们在VidOmni-Bench上的实验揭示了三个关键发现:(i)Video-LLMs在密集视频字幕生成中频繁产生幻觉描述;(ii)它们作为验证器时表现不佳,无法可靠检测看似合理但错误的事件描述;(iii)模型弱点随视频复杂度和时长而变化,揭示了当前Video-LLMs中多样且模型特定的瓶颈。

英文摘要

While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching, where models may succeed through superficial cues and incomplete annotations. To this end, we introduce VidOmni-Bench, a benchmark that requires models to verify whether each event in dense video captions is supported by the video. VidOmni-Bench consists of 500 videos spanning five complexity types and diverse durations from 4 seconds to 90 minutes. After collecting videos along these axes, we use diverse Video-LLMs to generate dense captions and obtain human-verified sentence-level labels, where sentences containing incorrect events serve as hard negatives for evaluation. Our experiments on VidOmni-Bench reveal three key findings: (i) Video-LLMs frequently generate hallucinated descriptions in dense video captioning; (ii) they also struggle as verifiers, failing to reliably detect plausible but incorrect event descriptions; and (iii) model weaknesses vary across video complexity and duration, revealing diverse, model-specific bottlenecks in current Video-LLMs.

发表机构

  • Korea University(高丽大学)
  • Sungkyunkwan University(成均馆大学)
  • The University of Tokyo(东京大学)
  • OMRON SINIC X Corporation(欧姆龙SINIC X公司)

机构由 AI 辅助整理,请以论文原文为准。

↑