超越像素:从视频先验到4D世界
Beyond Pixels: From Video Priors to 4D Worlds
AI总结:
本文提出Latent-to-4D方法,利用共享VAE的视频模型去噪潜在变量作为接口实现4D生成,在两个数据集上的DINO-F1指标优于对比方法,且受人类评估者认可。
AI中文摘要:
4D生成是从文本或图像等条件合成动态3D场景的技术。现有方法要么用独立的4D模型重建生成的RGB视频,要么适配特定视频生成器直接预测几何。前者存在分布不匹配和误差传播问题,后者将4D预测绑定到特定生成器,当生成器或条件机制变化时可能需要重新训练。本文探究:共享变分自编码器(VAE)的视频模型的最终去噪潜在变量,是否可作为可复用接口用于显式4D预测。基于该洞察,本文提出直接潜在变量到4D生成方法,实例化为Latent-to-4D,该方法绕过RGB,通过将视频潜在变量与预训练4D解码器的标记网格对齐,并通过逐帧和全局时空注意力优化来实现。该方法在约1000个现有重建片段上训练,单个检查点可在同一VAE家族内的多个视频扩散转换器中不变迁移。在Text4D-200和I4D-200数据集上,Latent-to-4D在基于投影的DINO-F1指标上,分别超过匹配的同潜在Wan+4RC级联模型2.88至3.45和5.81个点,同时在几何、时间稳定性和整体质量上更受人类评估者青睐。
英文摘要:
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.