AI 中文总结
该研究提出无3D的视频重拍摄范式TARS,通过自监督学习结合数据缩放与文本-相机联合条件设置,实现稳健的相机与视角控制,在大运动下合成未见区域的表现优于现有方法。
AI 中文摘要
视频重拍摄旨在生成具有可控相机运动和视角的视频。现有方法要么依赖显式3D先验,受限于重建质量,在合成未见区域时表现较差;要么依赖具有不同相机轨迹的配对视频,而此类数据的稀缺性阻碍了泛化。我们通过文本驱动的语义视角规格重新探讨视频重拍摄,以实现对拍摄尺度、视角及第一/第三人称视角的控制。为此,我们提出TARS,一种无3D的视频重拍摄范式。时间步敏感性分析表明,相机运动主要在高噪声阶段建立,此时会形成粗糙的时空结构。基于该见解,我们引入自监督训练,无需配对重拍摄数据或3D重建即可学习相机动态和基础视觉表示。通过数据缩放及文本-相机联合条件设置,TARS支持稳健的相机与视角控制,可在大相机运动下合理合成源视图之外的区域,同时支持反向视角重拍摄和视角切换。大量实验表明,TARS相比现有方法提供更准确、时间一致性更好的相机控制。项目页面:this https URL
英文摘要
Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods. Project Page: https://ymlinfeng.github.io/TARS.github.io/
Comments8 pages, 5 figures