ViD-GPT:在视频扩散模型中引入GPT式自回归生成
ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models
- Huawei Cloud Computing(华为云计算)
- Nanyang Technological University(南洋理工大学)
- Finvolution Group(融智集团)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出ViD-GPT,通过将GPT式因果生成和帧作为提示机制引入视频扩散模型,并结合kv-cache加速推理,实现了高质量、时间一致的长时间视频生成。
AI中文摘要:
随着扩散模型的进步,当今的视频生成已经取得了令人印象深刻的质量。但生成时间上一致的长时间视频仍然具有挑战性。大多数视频扩散模型(VDMs)以自回归方式生成长时间视频,即基于前一个片段的最后几帧来生成后续片段。然而,现有方法都涉及双向计算,这限制了每个自回归步骤的感受上下文,并导致模型缺乏长期依赖性。受大语言模型(LLMs)的巨大成功启发,并遵循GPT(生成式预训练Transformer),我们将因果(即单向)生成引入VDMs,并使用过去的帧作为提示来生成未来帧。对于因果生成,我们在VDM中引入因果时间注意力,这迫使每个生成的帧依赖于其之前的帧。对于帧作为提示,我们通过沿时间轴将条件帧与噪声帧(待生成的帧)拼接起来注入条件帧。因此,我们提出了视频扩散GPT(ViD-GPT)。基于这两个关键设计,在每个自回归步骤中,它能够从由所有先前生成的帧拼接而成的提示帧中获取长期上下文。此外,我们将kv-cache机制引入VDMs,消除了重叠帧的冗余计算,显著提高了推理速度。大量实验表明,我们的ViD-GPT在长时间视频生成上在定量和定性方面均达到了最先进的性能。代码将在https://github.com/Dawn-LX/Causal-VideoGen上提供。
英文摘要:
With the advance of diffusion models, today's video generation has achieved impressive quality. But generating temporal consistent long videos is still challenging. A majority of video diffusion models (VDMs) generate long videos in an autoregressive manner, i.e., generating subsequent clips conditioned on last frames of previous clip. However, existing approaches all involve bidirectional computations, which restricts the receptive context of each autoregression step, and results in the model lacking long-term dependencies. Inspired from the huge success of large language models (LLMs) and following GPT (generative pre-trained transformer), we bring causal (i.e., unidirectional) generation into VDMs, and use past frames as prompt to generate future frames. For Causal Generation, we introduce causal temporal attention into VDM, which forces each generated frame to depend on its previous frames. For Frame as Prompt, we inject the conditional frames by concatenating them with noisy frames (frames to be generated) along the temporal axis. Consequently, we present Video Diffusion GPT (ViD-GPT). Based on the two key designs, in each autoregression step, it is able to acquire long-term context from prompting frames concatenated by all previously generated frames. Additionally, we bring the kv-cache mechanism to VDMs, which eliminates the redundant computation from overlapped frames, significantly boosting the inference speed. Extensive experiments demonstrate that our ViD-GPT achieves state-of-the-art performance both quantitatively and qualitatively on long video generation. Code will be available at https://github.com/Dawn-LX/Causal-VideoGen.