发表机构
The Chinese University of Hong Kong, Shenzhen; Huazhong University of Science and Technology; Shenzhen Loop Area Institute; University of Science and Technology of China(香港中文大学(深圳); 华中科技大学; 深圳河套学院; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视频生成模型潜在空间问题,提出VideoRAE,利用冻结视频基础编码器特征经1D自注意力投影仪压缩,支持多种潜在空间,通过多码本高维量化等实现强大重建,收敛快,验证了冻结VFM表示的有效性。
AI 中文摘要
视频生成模型通常依赖于由3D变分自动编码器(3D-VAE)学习的潜在空间。然而,传统的3D-VAE主要针对像素级重建进行优化,这可能会限制其潜在空间捕获的语义和时空结构。同时,诸如V-JEPA 2和VideoMAEv2等视频基础模型(VFM)显示出强大的视频理解能力,但其冻结表示能否转换为紧凑、具有重建能力且对生成友好的视频潜在空间在很大程度上仍未得到探索。我们通过VideoRAE回答了这个问题,它是一种表示自动编码器,利用冻结视频基础编码器的多尺度分层特征,并通过轻量级1D自注意力投影仪对其进行压缩。VideoRAE通过多码本高维量化支持扩散变压器的连续潜在空间和自回归模型的离散令牌。在解码过程中,表示自动编码器通过冻结的VFM教师改进语义保留并实现无KL正则化的训练。实验表明,VideoRAE在连续和离散模式下均实现了强大的重建。在UCF-101上,它分别使用AR和DiT生成器获得了40和93的最新类到视频gFVD,同时收敛速度比竞争的自动编码器基线快约5倍。在受控的2B规模文本到视频研究中,在可比设置下,用VideoRAE替换LTX-VAE会导致更快的收敛。这些结果验证了冻结的VFM表示作为通用且对生成友好的视频潜在空间。模型和代码将在这个https URL上发布。
英文摘要
Video generation models typically rely on 3D-VAEs trained for pixel-level reconstruction, whose latent spaces may underrepresent semantic structure. We introduce VideoRAE, a representation autoencoder that converts features from a frozen video foundation model into compact, reconstruction-capable latents for video generation. A lightweight 1D self-attention projector compresses multi-scale hierarchical features, producing continuous latents for Diffusion Transformers and discrete tokens for autoregressive models through multi-codebook high-dimensional quantization. During decoding, a local-global representation alignment objective transfers semantic structure from the frozen encoder and removes the need for KL regularization. Comprehensive experiments show that VideoRAE achieves strong reconstruction in both continuous and discrete regimes. On UCF-101, autoregressive and diffusion generators built on VideoRAE achieve class-conditional gFVD scores of 40 and 93, respectively, while converging approximately five times faster than autoencoder baselines. In controlled 2B-parameter text-to-video experiments, replacing LTX-VAE with VideoRAE accelerates convergence and consistently improves VBench performance. These results establish frozen video foundation representations as compact, versatile, and generation-friendly video latents. Code and models are available at https://zhxie0117.github.io/VideoRAE/.
CommentsHome page: https://zhxie0117.github.io/VideoRAE