arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VidaForge:视频预训练数据配方的开放研究基础设施

VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

Yan Ma, Jiadi Su, Zhulin Hu, Ethan Chern, Linhao Zhang, Tiantian Mi, Pengfei Liu

arXiv 2609.06652首次发表:更新:

发表机构

Fudan University; Shanghai Jiao Tong University; Shanghai University; Shanghai Innovation Institute(复旦大学; 上海交通大学; 上海大学; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VidaForge提出一个开放的五阶段视频数据配方工作流,通过比较不同覆盖范围和质量的数据配方在Wan 2.1和V-JEPA 2.1预训练中的效果,证明其能连接数据配方选择与下游模型性能,并发布包含314万场景级片段、总计6475小时的VIDAFORGE-3M数据集。

AI 中文摘要

视频基础模型越来越依赖于大规模预训练数据,然而其背后的端到端数据流水线在很大程度上仍然封闭,难以检查或复用。研究人员若想理解视频数据配方如何影响模型预训练,往往需要在测试一个聚焦的假设之前就构建大量的基础设施。我们提出了VIDAFORGE,一个开放的研究基础设施,它将视频数据配方表示为一个从原始视频到训练数据集的可执行的五阶段工作流。该工作流中的某个决策可以被改变以构建替代数据集,同时保留每个样本是如何产生的。为了演示这一研究工作流,我们在Wan 2.1和V-JEPA 2.1的早期从头预训练中比较了具有不同覆盖范围和质量的数据配方。在这两种学习目标下,覆盖范围更广的配方取得了最高的下游基准分数,而基于损失的评估则偏好不同的配方。这项研究展示了VidaForge如何将数据配方的选择与下游模型性能联系起来。我们进一步发布了VIDAFORGE-3M,其中包含314万个场景级片段,总计6475小时,并带有细粒度的标注和用于视频数据配方研究的策展信号。

英文摘要

Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Researchers seeking to understand how video data recipes affect model pretraining often need to build substantial infrastructure before testing even a focused hypothesis. We present VIDAFORGE, an open research infrastructure that represents a video data recipe as an executable five-stage workflow from raw videos to training datasets. A decision in this workflow can be varied to construct alternative datasets while preserving how every sample was produced. To demon strate this research workflow, we compare data recipes with different coverage and quality in early from-scratch pretraining of Wan 2.1 and V-JEPA 2.1. Across both learning objectives, the broader-coverage recipe achieves the highest downstream benchmark scores, while loss-based evaluation favors different recipes. This study demonstrates how VidaForge connects data-recipe choices to downstream model performance. We further release VIDAFORGE-3M, containing 3.14 million scene level clips totaling 6,475 hours, with fine-grained annotations and curation signals for video data-recipe research.

Commentshttps://github.com/GAIR-NLP/VidaForge

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑