发表机构
Tsinghua University; Bosch (China) Investment Ltd.; University of Chinese Academy of Sciences(清华大学; 博世(中国)投资有限公司; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SymRegFlow提出对称正则化流匹配框架,通过几何扭曲与双锚点监督实现连续视点多视图一致视频生成,无需新视图RGB监督,在nuScenes上FVD降低超31%。
AI 中文摘要
基于流匹配的多视图世界模型能够生成逼真的视频,但通常局限于固定的相机装置。将它们扩展到连续变化的相机姿态,需要具有密集姿态覆盖的成对姿态-视频观测数据,而这类数据的获取成本高昂。我们提出了SymRegFlow,一个对称正则化的流匹配框架,用于在连续视点下生成多视图一致的视频,而无需真实新视图的RGB监督。对于每个目标姿态,SymRegFlow通过几何方式将源视图扭曲为带噪声的锚点,并结合掩蔽双锚点监督与跨锚点去噪输出一致性,以减轻锚点特定误差。在仿射高斯替代模型下,我们证明了合适的一致性正则化能够在固定噪声水平下恢复干净参考的最优解,严格优于单锚点和合并锚点的基线方法。在Cosmos-Drive-Dreams和nuScenes上的实验展示了高质量、多视图一致的自驾视频生成:在nuScenes上,SymRegFlow在所评估的基线中取得了最低的FVD和FVMD,相对于最佳基线将FVD降低了超过31%,并且基于源条件的推理还取得了最佳的FID和实例保持性能。
英文摘要
Flow-matching-based multi-view world models generate realistic videos, but are commonly restricted to fixed camera rigs. Extending them to continuously varying camera poses requires paired pose--video observations with dense pose coverage, which are costly to acquire. We introduce \emph{SymRegFlow}, a symmetry-regularized flow-matching framework for multi-view-consistent video generation across continuous viewpoints without ground-truth novel-view RGB supervision. For each target pose, SymRegFlow geometrically warps source views into noisy anchors and combines masked dual-anchor supervision with cross-anchor denoising-output consistency to mitigate anchor-specific errors. Under an affine Gaussian surrogate, we prove that suitable consistency regularization recovers the clean-reference optimum at fixed noise levels, strictly outperforming single- and merged-anchor baselines. Experiments on Cosmos-Drive-Dreams and nuScenes demonstrate high-quality, multi-view-consistent autonomous-driving video generation: on nuScenes, SymRegFlow achieves the lowest FVD and FVMD among the evaluated baselines, reducing FVD by over 31\% relative to the best baseline, and source-conditioned inference also attains the best FID and instance preservation.