arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06948cs.CV

PRG-Fusion:以重建证据编排生成先验用于驾驶视角合成

PRG-Fusion: Orchestrating Generative Priors with Reconstruction Evidence for Driving View Synthesis

Sipeng He, Jialei Chen, Zhen Fang, Dongchun Ren, Feng Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

提出PRG-Fusion框架,利用重建证据生成区域标签,编排3DGS保留、LiDAR修复和视频先验生成,实现高质量驾驶视角合成。

中文摘要 AI 辅助

沿指定轨迹合成逼真的驾驶视频对于可扩展的闭环仿真至关重要。基于重建的方法利用神经渲染来合成几何一致的视图,但当视点偏离训练轨迹时,往往会出现各种伪影和内容缺失。相比之下,生成模型可以从车辆传感器数据沿任意轨迹合成逼真的视图,但往往难以在帧间保持时间和几何一致性。为了结合两者的优势,我们提出了PRG-Fusion,一个用于驾驶视角合成的框架,它利用重建证据在区域间编排生成先验。具体来说,我们从重建的驾驶场景中提取逐区域的退化证据,并将其转换为保留(Preserve)、修复(Repair)和生成(Generate)标签。在推理时,这些标签作为统一的区域感知时空合成路由策略,分别编排3DGS外观保留、LiDAR引导的结构校正和视频先验驱动的内容补全,作用于保留、修复和生成区域。随后,我们采用两阶段训练范式,首先从稀疏LiDAR投影中建立几何控制,然后从密集3DGS渲染中学习外观控制。在Waymo上的大量实验表明,PRG-Fusion在新颖轨迹视频合成中达到了最先进的整体性能,具有优越的视觉质量和几何保真度,同时在大轨迹偏移下保持了有竞争力的视图一致性。

英文摘要

Synthesizing photorealistic driving videos along specified trajectories is essential for scalable closed-loop simulation. Reconstruction-based methods leverage neural rendering to synthesize geometrically consistent views, but often exhibit diverse artifacts and missing content when the viewpoint deviates from the training trajectory. In contrast, generative models can synthesize realistic views along arbitrary trajectories from vehicle sensor data, yet often struggle to maintain temporal and geometric consistency across frames. To combine the strengths of both, we propose PRG-Fusion, a framework for driving view synthesis that uses reconstruction evidence to orchestrate generative priors across regions. Specifically, we extract region-wise degradation evidence from reconstructed driving scenes and convert it into Preserve, Repair, and Generate (PRG) labels. At inference, these labels serve as a unified routing policy for region-aware spatiotemporal synthesis, orchestrating 3DGS appearance preservation, LiDAR-guided structural correction, and video-prior-driven content completion across Preserve, Repair, and Generate regions, respectively. We then follow a two-stage training paradigm, first establish geometric control from sparse LiDAR projections and subsequently learning appearance control from dense 3DGS renderings. Extensive experiments on Waymo demonstrate that PRG-Fusion achieves state-of-the-art overall performance in novel trajectory video synthesis, with superior visual quality and geometric fidelity while maintaining competitive view consistency under large trajectory shifts.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Yootta

机构由 AI 辅助整理,请以论文原文为准。

↑