推理时稳定相机控制的新视图合成
Stabilizing Camera-Controlled Novel View Synthesis at Inference Time
- IIT Gandhinagar(印度理工学院甘地纳格尔分校)
- Osaka University(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对预训练视频扩散模型的相机控制新视图合成在大运动下不稳定的问题,提出将相机运动分解为小自回归步骤等改进,使CamTrol++在多个数据集上提升了合成稳定性与效率。
AI中文摘要:
使用预训练视频扩散模型从单张图像进行无需训练、相机控制的新视图合成,在大相机运动和长生成范围下常出现不稳定问题。现有方法通常组合多个推理时组件,导致难以明确哪些设计选择对稳定性最重要。本文表明稳定性的主要来源很简单:将相机运动分解为小的自回归步骤,可限制每步几何失真并减少误差累积。受控相机步研究显示,小运动下性能保持稳定,每步运动接近18-20°时性能下降更明显。本文进一步评估几何约束空间注意力和低频外观锚定作为辅助改进,结合高效的无配准变形流水线。在RealEstate10K和MegaScene数据集上,CamTrol++相比无需训练的基线方法,在时间与几何一致性、下游3D重建质量及生成效率上均有提升,该方法在56帧生成及大量受控深度损坏下仍有效。这些结果表明,在推理时仔细控制相机运动,无需重新训练或修改扩散骨干网络,即可大幅提升相机控制的新视图合成的稳定性。
英文摘要:
Training-free, camera-controlled novel view synthesis from a single image using pre-trained video diffusion models often becomes unstable under large camera motion and long generation horizons. Existing approaches commonly combine several inference-time components, making it unclear which design choices are most important for stability. We show that the main source of stability is simple. Decomposing camera motion into small autoregressive steps limits per-step geometric distortion and reduces error accumulation. A controlled camera-step study shows that performance remains stable for small motions and degrades more strongly as the per-step motion approaches $18$-$20^\circ$. We further evaluate geometry-constrained spatial attention and low-frequency appearance anchoring as supporting refinements, together with an efficient registration-free warping pipeline. Across RealEstate10K and MegaScene, CamTrol++ improves temporal and geometric consistency, downstream 3D reconstruction quality, and generation efficiency over training-free baselines. The method remains effective for 56-frame generation and under substantial controlled depth corruption. These results show that careful control of camera motion at inference time can substantially improve the stability of camera-controlled novel view synthesis without retraining or modifying the diffusion backbone.