视频生成模型是通用视觉学习者
Video Generation Models are General-Purpose Vision Learners
浏览论文内容
中文总结 AI 辅助
研究探讨计算机视觉中实现通用模型的催化剂,提出大规模文本到视频生成是预训练范式。介绍GenCeption模型,利用视频生成扩散主干定义感知模型。实验表明该模型在多任务中性能领先,有数据效率优势且具涌现行为,为通用视觉智能提供基础路径。
中文摘要 AI 辅助
受下一个token预测驱动,自然语言处理从特定任务模型转变为强大的通用基础模型。那么,在计算机视觉中实现通用模型需要什么等效催化剂呢?本文认为大规模文本到视频生成是计算机视觉的强大预训练范式,提供通用视觉智能所需的时空先验、视觉语言对齐和可扩展性。我们引入GenCeption,它利用预训练的视频生成扩散主干来定义前馈感知模型,能够执行由文本指令引导的各种视觉任务。实证结果表明,GenCeption在各种任务中取得了领先性能,包括深度、表面法线和相机姿态估计、表情参考分割和3D关键点预测等,常常匹配或超越专门模型。视频生成预训练主干在可比设置下也优于其他预训练范式。GenCeption还展示了初步的数据和模型缩放属性以及卓越的数据效率,在训练数据少得多的情况下与领先模型取得可比性能。此外,GenCeption还表现出有趣的涌现行为:仅在合成人类视频上训练的模型能推广到真实世界镜头和分布外对象类别。这些发现表明视频生成不仅是一种合成工具,更是通往物理世界通用视觉智能的基础路径。
英文摘要
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world. Project page: https://genception.github.io
发表机构
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。