arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35090cs.CV

推进视频-文本预训练的多视角字幕

Advancing Video-Text Pretraining with Multi-View Captions

Fida M. Thoker, Renaud Vandeghen, Karen Sanchez, Marc Van Droogenbroeck, Bernard Ghanem

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种基于多模态大语言模型的多视角字幕生成框架,通过互补摘要与详细字幕、推理细化及语义正样本生成,提升视频-文本预训练中语言监督的质量,并在多个检索基准上以更小语料实现更优性能。

中文摘要 AI 辅助

视频-文本预训练通过模型和数据集的扩展取得了显著进展,然而语言监督的质量仍未得到充分探索。现有的网络规模数据集通常仅为每个视频提供单一稀疏字幕,无法捕捉丰富的时空语义,而直接使用字幕生成模型可能会产生嘈杂的描述。我们提出了一种基于大规模多模态大语言模型的监督生成框架,以提高监督的多样性、保真度和语义覆盖。从1000万个视频开始,我们的方法通过互补的摘要和详细字幕、基于推理的细化以及语义正样本字幕生成来产生多视角字幕(MVC)。为了有效利用不同粒度的监督,我们进一步引入了一种粒度感知的文本表示,为摘要和详细视图分别设置独立的CLS标记。我们使用生成的监督语料库预训练视频-文本模型,并在标准、细粒度和详细的文本到视频检索基准上进行评估。我们的方法在一致提升零样本和微调性能的同时,使用的预训练语料库比现有方法更小,展示了丰富且互补的文本监督对视频-文本预训练的重要性。项目页面:此https URL

英文摘要

Video-text pretraining has achieved remarkable progress through the scaling of models and datasets, yet the quality of language supervision remains underexplored. Existing web-scale datasets often provide only a single sparse caption per video that fails to capture rich spatiotemporal semantics, while directly using captioning models can generate noisy descriptions. We propose a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage. Starting from 10 million videos, our approach generates multi-view captions (MVC) through complementary summary and detailed captions, reasoning-based refinement, and semantic positive caption generation. To effectively exploit supervision at different granularities, we further introduce a granularity-aware text representation with separate CLS tokens for summary and detailed views. We pretrain video-text models using the resulting supervision corpus and evaluate them across standard, fine-grained and detailed text-to-video retrieval benchmarks. Our approach consistently improves both zero-shot and fine-tuned performance while using smaller pretraining corpora than existing methods, demonstrating the importance of rich and complementary textual supervision for video-text pretraining. Project page: https://rvandeghen.github.io/mvc/

发表机构

  • King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
  • University of Liège(列日大学)

机构由 AI 辅助整理,请以论文原文为准。

↑