视频模型作为原生4D渲染器:基于动画网格的世界 grounding 条件
Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
浏览论文内容
中文总结 AI 辅助
该研究提出参考引导渲染器DAR,扩展Wan2.2相机控制,用跟踪和世界位置组成的4D G缓冲作为条件,在DAR-4D基准上实现优于现有方法的视频生成性能。
中文摘要 AI 辅助
当期望的场景状态已由动画网格、相机轨迹和参考图像指定时,预训练的视频扩散模型可充当渲染器。这种4D生成渲染设置提出了一个表示问题:何种图像格式条件能让视频主干网络同时服从相机运动和场景内部动画?我们提出DAR,一种参考引导渲染器,它将Wan2.2的相机控制从仅基于Plücker射线扩展到相机加几何的联合接口。DAR从动画网格中投影神经4D G缓冲(跟踪、世界位置和法向量),并通过加宽的控制适配器注入,同时保留预训练的图像到视频先验。核心设计选择是跟踪和世界位置这一对元素:跟踪识别应承载外观的持久表面元素,世界位置提供其当前场景坐标状态,法向量提供局部形状。深度加校准射线原则上可恢复3D,但深度是依赖相机的图表,其中相机和物体运动被混合。在含68个案例的DAR-4D基准上,LoRA DAR达到PSNR 23.22、SSIM 0.895、LPIPS 0.134,比现成的Wan2.2-Depth提升1.54 dB PSNR;全微调则达到PSNR 25.36和SSIM 0.917。匹配的 ablation 实验显示,在每个检查点用深度替换世界位置会使PSNR降低1.26至1.55 dB,支持跟踪加世界位置对应关系作为实用的4D渲染条件。
英文摘要
Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D generative rendering setting raises a representation question: what image-format condition lets a video backbone obey both camera motion and scene-internal animation? We propose DAR, a reference-guided renderer that extends Wan2.2 camera control from Plücker rays alone to a joint camera-plus-geometry interface. DAR projects a neural 4D G-buffer (tracking, world position, and normal) from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior. The central design choice is the pair of tracking and world position. Tracking identifies the persistent surface element that should carry appearance; world position gives its current scene-coordinate state; normal supplies local shape. Depth plus calibrated rays can recover 3D in principle, but depth is a camera-dependent chart in which camera and object motion are mixed. On the 68-case DAR-4D benchmark, LoRA DAR reaches PSNR 23.22, SSIM 0.895, and LPIPS 0.134, improving over off-the-shelf Wan2.2-Depth by 1.54 dB PSNR; a full fine-tune reaches PSNR 25.36 and SSIM 0.917. Matched ablations show that replacing world position by depth reduces PSNR by 1.26--1.55 dB at every checkpoint, supporting tracking+world-position correspondence as a practical 4D rendering condition.
发表机构
- Tsinghua University(清华大学)
- The Hong Kong Polytechnic University(香港理工大学)
- University of Electronic Science and Technology of China(电子科技大学)
- Sun Yat-sen University(中山大学)
- The University of Hong Kong(香港大学)
- University of Science and Technology of China(中国科学技术大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。