移动字母表:文本到视频生成训练数据的对照研究
Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
查看机构详情
- Meta Superintelligence Labs(Meta超智能实验室)
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究文本到视频生成中数据分布和字幕质量对模型的影响,通过移动字母表测试平台进行对照实验,发现视频内容分布、字幕质量很关键,无分类器引导等可部分恢复模型,为相关模型开发提供参考。
中文摘要 AI 辅助
在过去五年中,文本到视频生成通过模型规模、数据和计算能力的扩展取得了显著进展。与模型架构不同,训练数据常常未得到充分探索。现实世界的数据管理复杂且具有挑战性,涉及从原始视频中选择片段并添加字幕以创建用于学习文本到视频映射的视频 - 文本对。我们研究数据分布和字幕质量如何影响文本到视频模型。为了进行对照实验,我们引入了移动字母表,这是一个程序测试平台,可在黑色背景上以不同字体、颜色、大小和位置呈现字母,并以不同方向和速度移动。这种设计允许通过破坏真实元数据来精确控制数据分布和字幕质量。我们的实验得出了三个发现:a)视频内容和时长的多样化且平衡的分布对于泛化至关重要;b)字幕质量显著影响模型性能和训练效率,这表明文本到视频模型受视频理解能力的限制;c)无分类器引导和在高质量数据上的微调可以部分恢复在损坏字幕上训练的模型,但无法完全弥补预训练数据不佳的问题。我们相信这些见解可以为大规模文本到视频模型的开发提供参考,并且我们主张更加关注预训练数据的科学。
英文摘要
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.