arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

4DStreamCtrl:基于在线4D控制的交互式视频生成

4DStreamCtrl: Interactive Video Generation with Online 4D Control

Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu

arXiv 2608.25479首次发表:更新:

发表机构

Peking University; Tencent Hunyuan(北京大学; 腾讯混元)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

4DStreamCtrl将相机运动、物体轨迹与深度统一为3D点轨迹表示,结合蒸馏的因果流式模型,首次实现交互式4D可控实时视频生成,精度与效率均优于现有方法。

AI 中文摘要

生成式视频模型如今可合成与现实几乎难以区分的画面,其作为交互式工具的潜力取决于对物体和相机随时间运动的细粒度控制,但现有方法均仅能实现部分控制:相机参数方法可操控视角但无法移动物体,2D轨迹方法在图像平面内作用且忽略深度与遮挡,近期的3D方法虽加入几何信息却仅能离线生成固定长度视频。尤其,尚无方法能同时实现相机与物体的3D一致性控制及实时流式生成。本文提出将相机运动、物体轨迹与深度统一为单一3D点轨迹表示,基于该表示,单个模型可在单次前向传播中实现相机与物体联合控制、深度编辑及运动迁移。为大规模学习该接口,本文从野外视频中挖掘3D运动监督数据,构建OpenVidHD-Motion3D数据集,并通过轻量型几何运动头(Geometric Motion Head)对其编码,该几何运动头可接入预训练视频扩散模型。由于该编码器具备时间可分性,本文将模型蒸馏为因果流式学生模型,该模型可生成任意长度视频,仅需4次去噪步骤,且内存占用与视频长度无关。该统一设计在运动控制精度上超越了仅控制相机、2D及离线3D方法,同时覆盖了这些方法仅单独处理的模态。4DStreamCtrl在单块高端GPU上以20 FPS帧率生成480p视频,且在数百帧内保持时间一致性,据本文所知,其首次实现了交互式4D可控流式生成。更广泛而言,将生成过程基于显式3D几何并结合高效因果推理,有望催生具备闭环时空控制的交互式世界模型,从可控模拟器到具身智能体的实时视觉想象。

英文摘要

Generative video models now synthesize footage nearly indistinguishable from reality. Their promise as interactive tools hinges on fine-grained control of how objects and the camera move over time, yet each existing approach captures only part of this: camera-parameter methods steer the viewpoint but cannot move objects, 2D-trajectory methods act in the image plane and ignore depth and occlusion, and recent 3D methods add geometry but run only offline at a fixed length. In particular, none combines 3D-consistent control of both camera and objects with real-time, streaming generation. Here we show that camera motion, object trajectories, and depth can be unified into a single 3D point-track representation, from which one model performs joint camera and object control, depth editing, and motion transfer in a single forward pass. To learn this interface at scale, we mine in-the-wild video for 3D motion supervision, yielding OpenVidHD-Motion3D, and encode it with a lightweight Geometric Motion Head that plugs into a pretrained video diffusion model. Because this encoder is temporally separable, we distill the model into a causal streaming student that generates arbitrarily long video in four denoising steps at memory independent of length. This unified design surpasses prior camera-only, 2D, and offline-3D methods in motion-control precision while covering modalities they address only in isolation. 4DStreamCtrl runs at 20 FPS on a single high-end GPU for 480p video and stays temporally coherent over hundreds of frames, enabling, to our knowledge, interactive 4D-controllable streaming generation for the first time. More broadly, grounding generation in explicit 3D geometry with efficient causal inference points toward interactive world models with closed-loop spatiotemporal control, from controllable simulators to real-time visual imagination for embodied agents.

Comments23 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑