arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式电影摄影师:在3D中编排相机与物体运动

Generative Cinematographer: Composing Camera and Object Motion in 3D

Jiahan Zhang, Chaohao Yang, Namitha Guruprasad, Vivekjyoti Banerjee, Trong-Tung Nguyen, Alan Yuille, Anand Bhattad

arXiv 2610.02180首次发表:更新:

发表机构

Johns Hopkins University(约翰霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有视频生成中2D控制歧义问题,提出GenCine系统,将单图提升为3D场景脚手架,通过局部3D手柄联合编排相机与前景运动,并投影为引导图控制预训练模型,实现相机相对运动一致与强可控性。

AI 中文摘要

当前可控视频生成系统通常依赖2D运动轨迹或稀疏拖拽信号来控制物体运动。这些控制具有歧义性,因为相同的2D轨迹可能对应不同的3D运动,尤其是当相机和物体同时移动时。我们提出了生成式电影摄影师(GenCine),该系统将单张图像提升为可编辑的3D场景脚手架,艺术家可在其中共同编排相机和前景运动。艺术家指定相机路径,并使用局部3D运动手柄移动选定的前景区域。多个手柄可以独立移动主体的不同部分,从而在无需物理模拟器或类别特定先验的情况下,提供对非刚性运动的分段刚性近似。为了将这些控制传达给预训练的视频模型,我们将其投影为引导图。这些图记录受控区域在每一帧中出现的位置,为每个手柄分配跨帧固定的颜色,并在与背景相同的世界坐标系中编码其受控点的当前3D位置。这使得我们即使在相机移动时也能描述物体相对于场景的运动。在训练方面,我们从真实视频中观察到的运动恢复控制,并使用合成视频中的真实几何和轨迹。我们在预训练的Wan模型上训练了一个轻量级引导分支和LoRA适配器,以遵循这些控制。我们的实验表明,该方法在视角变化下具有一致的相机相对运动、改进的几何一致性,并在多样化的真实世界场景中展现出强大的可控性。

英文摘要

Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑