arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28174cs.CV

Manifold4D:在点云渲染流形上进行去噪以实现视频重拍摄

Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting

Yongqi Mao, Zijia Dai, Zhishuo Liu, Wei Xu, Kaiwei Wang, Guotao Meng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出MANIFOLD4D,将渲染结果注入流匹配初始噪声以实现视频重拍摄,在DAVIS-Traj等基准上显著提升相机控制精度,同时保证视频保真度与动态一致性。

中文摘要 AI 辅助

视频重拍摄是指沿着用户指定的相机轨迹对动态场景的单目视频进行重新渲染,主流方案会显式提供目标几何结构:通过逐帧深度将源视频提升为4D点云,再沿轨迹光栅化为点云渲染结果。由于渲染结果和源视频均作为视觉条件输入网络,二者在每一步去噪过程中会产生竞争,使模型面临信任困境——即需要确定对渲染结果的信任程度,这可能会降低训练分布外数据的轨迹控制或视觉质量。本文认为,已与目标视图像素对齐的渲染结果根本不需要作为显式条件流提供。我们提出MANIFOLD4D,该方法将渲染结果直接注入流匹配的初始噪声中,使生成过程不再从标准高斯噪声开始,而是从携带几何信息的新噪声流形开始,仅将源视频作为唯一视觉条件。渲染结果仅使用一次,网络无需学习如何读取它;在后续去噪步骤中,模型可专注于源视频。在DAVIS-Traj基准和Vista4D评估集上,MANIFOLD4D在所有指标上均实现了最佳相机控制精度,与最强基线相比,旋转误差分别降低25%和27%,平移误差最高降低32%,同时在视频保真度上与基线相当,并在真实世界新视图光度质量上领先。用户研究表明,本文方法在轨迹跟踪和动态一致性上具有明显优势;当偏航幅度超出训练范围时,优势会进一步扩大,且当渲染结果被故意破坏时,模型仍能从源视频中恢复正确的动态运动,证实几何先验在引导生成的同时不会覆盖生成过程。

英文摘要

Video re-shooting re-renders a monocular video of a dynamic scene along a user-specified camera trajectory, and the dominant recipe supplies the target geometry explicitly: per-frame depth lifts the source video into a 4D point cloud, which is rasterized along the trajectory into a point cloud render. Because the render and the source video are both handed to the network as visual conditions, they compete at every denoising step, leaving the model with a trust dilemma --- how much of the render to believe --- which can degrade trajectory control or visual quality on data outside the training distribution. We argue that a render already pixel-aligned with the target view does not need to be supplied as an explicit conditioning stream at all. We propose MANIFOLD4D, which injects the render directly into the initial noise of flow matching, so that generation no longer departs from standard Gaussian noise but from a new noise manifold carrying geometric information, leaving the source video as the only visual condition. The render is thus used exactly once, and the network is never asked to learn how to read it; in subsequent denoising steps the model can focus on the source video. On our DAVIS-Traj benchmark and on the Vista4D evaluation set, MANIFOLD4D attains the best camera-control accuracy on every metric, lowering rotation error by 25% and 27% and translation error by up to 32% over the strongest baseline, while matching it in video fidelity and leading on real-world novel-view photometric quality. In a user study, our method achieves clear advantages in trajectory following and dynamic consistency. The gap widens as the yaw amplitude grows past the training range, and the model still recovers correct dynamic motion from the source video when the render is deliberately corrupted, confirming that the geometric prior guides generation without overriding it.

发表机构

  • Zhejiang University(浙江大学)
  • Manifold Tech(曼尼福尔德科技)
  • ShanghaiTech University(上海科技大学)
  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑