PlaylistEval:视频语言裁判能否在日级及更长时间尺度上被信任?
PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出PlaylistEval框架,基于100小时播放列表自动构建视频语言裁判基准,无需人工标注,生成630对需跨集检索的问题,评估17个模型显示前沿裁判准确率仅75.4%,且多模态和集合规模影响判断可靠性。
AI中文摘要:
视频语言模型越来越多地被用作视频理解的裁判,既用于评估模型输出,也用于训练奖励模型。当证据埋藏在长达一天的视频中时,它们的判断是否仍然可靠尚未得到证实。现有基准无法回答这个问题。它们的视频通常只有几分钟长,许多答案对可以仅从转录文本中分离出来,并且收集人工判断无法扩展到超长视频。我们引入了PlaylistEval,一个智能体框架,它可以在无需人工标注的情况下,基于100小时的播放列表集合构建视频语言裁判基准。它自动生成带有配对答案的问题,这些答案的差异由因果退化控制,因此每一对都需要跨集合进行检索。由此产生的基准包含跨越七个领域的630对,涵盖静态和动态知识,在一个由152对组成的分层子集上,它与人工判断的一致性达到93.0%(IAA 0.781)。评估来自八个家族的17个全模态和多模态模型显示,前沿裁判仅达到75.4%的成对准确率,而开源裁判模型的表现则远远落后。我们进一步表明,检索和最终判断都依赖于使用多种模态,并且随着播放列表集合的增长,裁判准确率会下降。我们在以下网址发布我们的流程、基准和评估代码:此https URL。
英文摘要:
Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at https://playlisteval.github.io.