VideoGen:一种用于高清文本到视频生成的参考引导潜在扩散方法
VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation
- Baidu Inc.(百度公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出VideoGen,利用文本到图像模型生成的参考图像引导级联潜在扩散模块,结合时间上采样与增强视频解码器,实现高保真、强时间一致性的高清文本到视频生成,达到SOTA。
AI中文摘要:
在本文中,我们提出了VideoGen,一种文本到视频生成方法,该方法使用参考引导的潜在扩散,能够生成具有高帧保真度和强时间一致性的高清视频。我们利用现成的文本到图像生成模型(例如Stable Diffusion),从文本提示生成具有高内容质量的图像,作为引导视频生成的参考图像。然后,我们引入一个高效的级联潜在扩散模块,该模块以参考图像和文本提示为条件,用于生成潜在视频表示,随后进行基于流的时间上采样步骤以提高时间分辨率。最后,我们通过增强的视频解码器将潜在视频表示映射为高清视频。在训练期间,我们使用真实视频的第一帧作为参考图像来训练级联潜在扩散模块。我们方法的主要特征包括:由文本到图像模型生成的参考图像提高了视觉保真度;将其作为条件使扩散模型更专注于学习视频动态;并且视频解码器在无标签视频数据上进行训练,从而受益于高质量且易于获取的视频。VideoGen在定性和定量评估方面都为文本到视频生成设定了新的最先进水平。更多样本请参见\url{https://videogen.github.io/VideoGen/}。
英文摘要:
In this paper, we present VideoGen, a text-to-video generation approach, which can generate a high-definition video with high frame fidelity and strong temporal consistency using reference-guided latent diffusion. We leverage an off-the-shelf text-to-image generation model, e.g., Stable Diffusion, to generate an image with high content quality from the text prompt, as a reference image to guide video generation. Then, we introduce an efficient cascaded latent diffusion module conditioned on both the reference image and the text prompt, for generating latent video representations, followed by a flow-based temporal upsampling step to improve the temporal resolution. Finally, we map latent video representations into a high-definition video through an enhanced video decoder. During training, we use the first frame of a ground-truth video as the reference image for training the cascaded latent diffusion module. The main characterises of our approach include: the reference image generated by the text-to-image model improves the visual fidelity; using it as the condition makes the diffusion model focus more on learning the video dynamics; and the video decoder is trained over unlabeled video data, thus benefiting from high-quality easily-available videos. VideoGen sets a new state-of-the-art in text-to-video generation in terms of both qualitative and quantitative evaluation. See \url{https://videogen.github.io/VideoGen/} for more samples.