发表机构
The Chinese University of Hong Kong; The Hong Kong Polytechnic University; The University of New South Wales; Southeast University(香港中文大学; 香港理工大学; 新南威尔士大学; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Streaming4D提出分块视频生成与增量重建结合的同步流水线,在RTX 4090上实现1.24倍运行时加速,同时保持4D几何与多视图一致性,解决传统4D生成延迟高的问题。
AI 中文摘要
当前4D生成范式常受限于顺序解耦设计:先生成视频,再进行3D重建,导致交互延迟高,限制了其在交互式实时场景中的应用。为此,我们提出Streaming4D,一种紧密耦合的同步流水线,将分块自回归视频生成与增量3D重建相结合。与传统逐帧生成和延迟几何恢复不同,Streaming4D生成时间视频块,并在每个块完成后立即触发重建,使合成与几何更新并行执行。该方法让世界表示随视频流在线演化,在保留几何保真度的同时降低反馈延迟。我们采用自强迫式自回归生成器和增量重建后端实现Streaming4D。实验表明,在单张RTX 4090上,该方法在各分辨率下均实现一致的运行时提升(加速1.24倍),同时保持高质量的4D几何和多视图一致性。
英文摘要
Current 4D generation paradigms are often bottlenecked by a sequential decoupling design: video is generated first, followed by 3D reconstruction, leading to high interaction latency. This limits applications in interactive real-time scenarios. To this end, we propose \textbf{Streaming4D}, a tightly coupled synchronous pipeline that integrates block-wise autoregressive video generation with incremental 3D reconstruction. Unlike traditional frame-by-frame emission and delayed geometry recovery, Streaming4D generates temporal video blocks and immediately triggers reconstruction for each completed block, enabling parallel execution between synthesis and geometric updates. This approach allows the world representation to evolve online with the video stream, reducing feedback latency while preserving geometric fidelity. We instantiate \textbf{Streaming4D} using a Self-Forcing-style autoregressive generator and an incremental reconstruction backend. Experiments show consistent runtime improvements across resolutions on a single RTX 4090 (1.24$\times$ speedup), while maintaining high-quality 4D geometry and multi-view consistency.
CommentsAccepted by CVPR 2026 4DV Workshop