发表机构
University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出VideoTok4D,一种4D感知视频分词器,通过时空解耦、轨迹感知注意力和扩散先验,实现紧凑世界表示,性能最优且存储量降低4个数量级。
AI 中文摘要
视频分词器已成为现代视频建模的基石,通过将高维视觉信号映射到紧凑的潜在空间,支撑了压缩、重建和生成方面的进展。然而,尽管取得了这些进展,当前的分词范式在很大程度上仍停留在2D视觉领域,将视频视为图像序列,而非对底层动态3D世界的观测。因此,学习到的词元继承了这种以观测为中心的偏差,限制了其紧凑表示真实世界4D场景的能力。为缓解这一问题,我们提出了VideoTok4D,一种新颖的4D感知视频分词器,用于紧凑的世界表示。具体而言,我们的方法包含三个关键设计:1)一种时空解耦策略,将视频分解为静态和动态词元,以实现整体世界建模;2)一种轨迹感知的动态注意力机制,聚合轨迹对齐的线索以促进跨视角运动一致性;3)Co4DGen,一种在所得VideoTok4D词元空间上学习的扩散先验,用于高效的4D场景生成。大量实验表明,我们提出的方法实现了最先进的性能,同时所需存储量比密集4D表示最多低4个数量级。此外,紧凑的词元空间大幅缩短了扩散序列,从而实现了高效生成。
英文摘要
Video tokenizers have emerged as a cornerstone of modern video modeling, underpinning progress in compression, reconstruction and generation by mapping high-dimensional visual signals into compact latent spaces. However, despite this progress, current tokenization paradigms largely remain within the 2D visual domain, treating videos as image sequences rather than observations of an underlying dynamic 3D world. Consequently, the learned tokens inherit this observation-centric bias, limiting their capacity to compactly represent real-world 4D scenes. To mitigate this issue, we propose VideoTok4D, a novel 4D-aware video tokenizer for compact world representation. Specifically, our approach comprises three key designs: 1) a spatiotemporal disentanglement strategy that factorizes videos into static and dynamic tokens for holistic world modeling; 2) a track-aware dynamic attention mechanism that aggregates trajectory-aligned cues to promote cross-view motion consistency; and 3) Co4DGen, a diffusion prior learned over the resulting VideoTok4D token space for efficient 4D scene generation. Extensive experiments have demonstrated that our proposed method achieves state-of-the-art performance while requiring up to 4 orders of magnitude less storage than dense 4D representations. Moreover, the compact token space substantially shortens diffusion sequences, enabling efficient generation.
Comments9 pages, 5 figures