arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13556cs.CV

V-RAE:重新思考用于生成的视频隐空间

V-RAE: Rethinking Video Latent Spaces for Generation

  • National University of Singapore(新加坡国立大学)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Minghui Guo, Shengqiong Wu, Hao Fei

AI总结:

本研究提出V-RAE视频表示自动编码器,基于冻结视觉基础模型构建生成隐空间,在多项视频任务上优于现有模型,还引入tFVD指标,证明冻结语义表示可支持多种视频建模任务。

AI中文摘要:

隐式视频生成依赖自动编码器来定义生成模型运行的紧凑空间。尽管视频自动编码器架构已取得显著进展,但其隐空间仍主要针对像素级重建进行优化,仅提供有限的高层语义组织。然而,重建最优的隐空间未必适用于生成建模。我们提出V-RAE,这是一种视频表示自动编码器,它在冻结的视觉基础模型表示之上构建紧凑的生成隐变量。轻量级时间池化模块可消除时间冗余同时保留语义结构,视频解码器则从压缩特征中重建连续运动。我们在视频重建、语义探测和类条件生成任务上,使用四个代表性冻结编码器对V-RAE进行评估。V-RAE在K600数据集上达到2.13的rFVD,优于所有评估的大规模预训练视频VAE。其隐变量比传统视频分词器隐变量保留了多得多的语义信息。在匹配的生成设置下,我们的最佳变体在UCF101和K600上分别达到117.86和19.16的gFVD,同时收敛速度快达6倍。我们进一步表明,仅重建质量不足以表征生成效用,并引入tFVD,这是一种与下游生成质量更可靠相关的时间一致性诊断指标。除视频生成外,V-RAE在匹配的预测设置下,在Cityscapes数据集上的未来视频预测也优于Wan 2.2 VAE隐空间。总体而言,实验表明冻结的语义表示可支持视频重建、生成和预测建模。项目页面:this https URL。

英文摘要:

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations. A lightweight temporal pooling module removes temporal redundancy while preserving semantic structure, and a video decoder reconstructs continuous motion from the compressed features. We evaluate V-RAE with four representative frozen encoders on video reconstruction, semantic probing, and class-conditional generation. V-RAE achieves 2.13 rFVD on K600, outperforming all evaluated large-scale pretrained video VAEs. Its latents retain substantially more semantic information than conventional video tokenizer latents. Under matched generation settings, our best variant achieves gFVD scores of 117.86 and 19.16 on UCF101 and K600, respectively, while converging up to 6x faster}. We further show that reconstruction quality alone is insufficient to characterize generative utility and introduce tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality. Beyond video generation, V-RAE also improves future video prediction on Cityscapes over the Wan 2.2 VAE latent space under matched prediction settings. Taken together, the experiments show that frozen semantic representations can support video reconstruction, generation, and predictive modeling. The project page: https://v-rae.github.io/.

补充信息

相关深度报道

↑