发表机构
South Dakota State University(南达科他州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CLIP-CC-Bench是针对视频语言模型的长篇段落级视频描述评估套件,采用多LLM嵌入模型集成与粗细粒度语义匹配方法,评估17种模型填补了现有短片段基准的空白。
AI 中文摘要
视频语言模型的基准测试大多集中在短片段和单句指标上,目前尚不清楚现有系统是否能生成准确的长篇段落级描述。我们推出CLIP-CC-Bench,这是一个针对长篇视频描述的评估套件,由5小时电影内容分段为90秒的片段构建,每个片段配有专家撰写的段落式参考文本。该评估套件采用五个最先进的基于LLM的嵌入模型组成的集成以提高可靠性并减轻单模型偏差,应用两种互补方法:(i)粗粒度语义匹配和(ii)细粒度语义匹配,以比较模型生成的描述与CLIP-CC-Bench参考文本。使用该框架,我们评估了17个最先进的视频语言模型,报告了它们的Borda聚合排名和在CLIP-CC-Bench上的平均得分。我们还通过评判者间一致性和自举排名稳定性量化了该协议的内部可靠性。我们在此httpsURL发布了标准化评估脚本、模型输出和聚合工具以支持可复现性。CLIP-CC-Bench为长篇视频描述提供了实用的评估框架,填补了现有短片段和仅QA基准留下的空白。
英文摘要
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference. The evaluation suite employs an ensemble of five state-of-the-art LLM-based embedding models to increase reliability and mitigate single-model bias, and applies two complementary methodologies: (i) coarse-grained semantic matching and (ii) fine-grained semantic matching to compare model-generated descriptions against CLIP-CC-Bench references. Using this framework, we evaluate 17 state-of-the-art video-language models and report both their Borda-aggregated rankings and their average scores on CLIP-CC-Bench. We further quantify the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability. We release standardized evaluation scripts, model outputs, and aggregation tools at https://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Bench to support reproducibility. CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks.
CommentsAccepted and presented at EvalMG 2026, the Second Workshop on Evaluation for Multimodal Generation, co-located with ACM SIGIR 2026