AI 中文总结
研究团队构建了千万小时级开源多模态视频数据集LAION-BVD,经其训练的模型在视频-文本、音频-文本及图像-文本基准上表现出色,为多模态预训练提供了大规模数据支撑。
AI 中文摘要
我们提出LAION-BVD,这是一个面向多模态学习的大规模开源视频数据集,包含从CommonCrawl收集的13亿条平台专属视频URL。我们从这些URL中下载了8000万个视频,总时长达到1000万小时。该数据集专为视频、音频和图像模态的多模态预训练而设计。通过基于内容感知的场景检测,我们提取视频片段并为其合成生成视频和音频字幕。在这些数据上训练的模型在标准视频-文本和音频-文本基准上取得了有竞争力的性能,且随着训练量或模型规模的增加,性能会持续提升。此外,我们通过提取场景变化帧,探索将视频帧作为图像-文本数据的替代来源。这些帧呈现出与标准网络图像语料库不同的视觉分布,在该数据集上训练的模型实现了出色的图像-文本检索性能。我们将LAION-BVD发布给研究界,它以前所未有的规模显著扩大了多模态视频的开源获取渠道。
英文摘要
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.