DC-VideoGen:使用深度压缩视频自编码器的高效视频生成
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出 DC-VideoGen 训练后加速框架,结合深度压缩视频自编码器与 AE-Adapt-V,使 Wan-2.1-14B 以低成本适配高压缩潜空间,最高降低 14.8 倍延迟并支持单卡 4K 视频生成。
AI中文摘要:
我们提出 DC-VideoGen,这是一个用于高效视频生成的训练后加速框架。DC-VideoGen 可应用于任何预训练视频扩散模型,通过轻量级微调使其适配深度压缩潜空间,从而提升效率。该框架建立在两项关键创新之上:(i)一个深度压缩视频自编码器,采用新颖的块因果时间设计,实现 32 倍/64 倍空间压缩和 4 倍时间压缩,同时保持重建质量以及对更长视频的泛化能力;(ii)AE-Adapt-V,一种稳健的适配策略,可将预训练模型快速、稳定地迁移到新的潜空间。使用 DC-VideoGen 适配预训练 Wan-2.1-14B 模型在 NVIDIA H100 GPU 上仅需 10 个 GPU 日。加速后的模型推理延迟相比基础模型最高降低 14.8 倍,且不损失质量,并进一步支持在单张 GPU 上生成 2160×3840 视频。代码:https://github.com/dc-ai-projects/DC-VideoGen。
英文摘要:
We introduce DC-VideoGen, a post-training acceleration framework for efficient video generation. DC-VideoGen can be applied to any pre-trained video diffusion model, improving efficiency by adapting it to a deep compression latent space with lightweight fine-tuning. The framework builds on two key innovations: (i) a Deep Compression Video Autoencoder with a novel chunk-causal temporal design that achieves 32x/64x spatial and 4x temporal compression while preserving reconstruction quality and generalization to longer videos; and (ii) AE-Adapt-V, a robust adaptation strategy that enables rapid and stable transfer of pre-trained models into the new latent space. Adapting the pre-trained Wan-2.1-14B model with DC-VideoGen requires only 10 GPU days on the NVIDIA H100 GPU. The accelerated models achieve up to 14.8x lower inference latency than their base counterparts without compromising quality, and further enable 2160x3840 video generation on a single GPU. Code: https://github.com/dc-ai-projects/DC-VideoGen.