发表机构
Meta AI; Massachusetts Institute of Technology; KAUST; Duke University(Meta AI; 麻省理工学院; 阿卜杜拉国王科技大学; 杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探索用原始视频对Qwen3-1.7B进行中间训练,通过预测下一个视觉标记提升了模型在视频、图像基准测试的性能,且未损失文本性能,证明纯自监督视频中间训练的有效性。
AI 中文摘要
多模态大语言模型主要从配对的图像-文本数据或带注释的视频中学习,而原始网络视频很少被用于进一步训练现有语言模型。本研究探讨无字幕、无文本损失的原始视频是否可作为预训练语言模型的中间训练数据。研究将帧编码为连续的视觉标记,让语言模型学习预测下一个视觉标记。研究在YT-Temporal-1B的原始片段上对Qwen3-1.7B进行中间训练,随后对该模型和未进行中间训练的模型应用相同的图像-文本指令调优,使两者仅在中间训练环节存在差异。在涵盖感知、文档和图表任务的四个视频基准测试中,中间训练后的模型平均得分高出2.9分;在十个图像基准测试中,平均得分高出5.1分。尽管中间训练不包含文本,文本性能仍得以保留:在14个文本基准测试中,中间训练后模型的平均得分为48.9,而未进行中间训练的模型为48.0。训练过程分析显示,图像和视频性能提升在训练的30%阶段内出现,之后趋于稳定,波动幅度小于0.5分。此外,预测字幕的表现未能优于预测下一个视觉标记,表明视频中间训练可完全保持自监督模式,无需承担自动字幕带来的计算开销或标注噪声。
英文摘要
Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded into continuous visual tokens, and the language model learns to predict the next visual token. We mid-train Qwen3-1.7B on raw clips from YT-Temporal-1B and then apply the same image-text instruction tuning to it and to the model without mid-training, so that the two differ only in mid-training. The mid-trained model scores 2.9 points higher on average across four video benchmarks and 5.1 points higher across ten image benchmarks, spanning perception, document, and chart tasks. Text performance is preserved even though mid-training includes no text, with an average of 48.9 across 14 text benchmarks compared with 48.0 for the model without mid-training. Analyses across training show that the image and video gains emerge within 30% of training and plateau thereafter, varying by less than 0.5 points. Predicting captions fails to outperform next-visual-token prediction, demonstrating that video mid-training can remain purely self-supervised without the computational overhead or labeling noise of automated captioning.