arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02160cs.CV

4Director:利用刚性三维几何控制视频世界模型

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

首次发表
浏览论文内容

中文总结 AI 辅助

4Director通过显式四维场景表示和运动适配器,实现对视频中相机与物体运动的精确控制,并在新数据集上验证了其优越性。

中文摘要 AI 辅助

精确控制相机和物体运动对于专业视频制作至关重要。现有方法仅通过图像平面线索(在深度和旋转方面具有模糊性)或通过缺乏完整几何形状且在不同视角变化下失去一致性的三维轨迹和斑点,对物体进行粗略控制。我们引入了4Director,一种以显式四维场景表示为条件的视频世界模型:每个物体从输入图像中重建一次,作为规范网格,并通过每帧一个规定的刚性变换进行移动。这种表示提供了一种直观的三维控制界面,并防止未观察到的几何形状在每一帧中被独立重新生成。我们将受控场景渲染为深度视频,并引入了一个运动适配器,该适配器将此几何支架转换为视频,同时合成视角一致的表面外观、光照和非刚性动态。为了训练,我们构建了RealCOD-Rigid,这是一个由我们的自动管道标注的20,774个剪辑的新数据集,包含刚性三维场景。我们进一步引入了身份门控交并比(IG-IoU),该指标联合评估对规定物体运动的遵循程度和物体身份的保持。实验表明,4Director在视觉质量以及相机和物体控制方面始终优于先前方法。

英文摘要

Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.

发表机构

  • Stability AI
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑