arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将字幕置于测试之下:通过多项选择问答评估视频字幕质量

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Zizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li, Yang Lu, Meng Cao, Ping Huang, Xiaoming Simon Wang

arXiv 2609.09973首次发表:更新:

发表机构

Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有视频字幕评估指标的局限,提出无参考基准CapQuiz,通过人工验证的多项选择问答衡量信息保真度,实验证明其与人类判断相关性更高。

AI 中文摘要

评估视频字幕生成仍然是视觉大语言模型(VLLMs)面临的一个关键挑战。现有指标主要依赖于将生成的文本与地面真值参考进行匹配。这种范式受到视频描述“一对多”特性的困扰,高质量字幕常常因词汇不匹配或视觉焦点的合理转移而受到惩罚。此外,此类评估通常是单维度的,无法对字幕质量进行细粒度分析。为解决这一问题,我们通过信息保真度的视角重新定义字幕质量:字幕必须最大化对显著视觉信息的覆盖,同时确保严格的事实性。我们引入了CapQuiz,一种新颖的无参考基准,该基准根据字幕在回答源自视频的人工验证的、细粒度的多项选择问题中的效用性来评估字幕。CapQuiz具有一个包含10种问题类型(涵盖描述性和推断性类别)的层次化分类体系,覆盖24个多样的视频领域。大量实验表明,与现有指标相比,CapQuiz与人类判断的相关性显著更高,并为模型性能提供了可解释的见解。

英文摘要

Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the ``one-to-many'' nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality through the lens of information fidelity: A caption must maximize the coverage of salient visual information while ensuring strict factuality. We introduce CapQuiz, a novel reference-free benchmark that assesses captions based on their utility in answering human-verified, fine-grained, multiple-choice questions derived from the video. CapQuiz features a hierarchical taxonomy of 10 question types (spanning Descriptive and Inferential categories) across 24 diverse video domains. Extensive experiments demonstrate that CapQuiz correlates significantly better with human judgments than existing metrics and offers interpretable insights into model performance.

CommentsAccepted by ACL 2026 main conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑