arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPVC:面向跨数据集驾驶场景渲染的结构化全景视频修复

SPVC: Structured and Panoptic Video Fixing for Cross-Dataset Driving Scene Rendering

Gen Li, Shu Han, Yun Xi Qiao, Hua Chen, Xuyang Dai, Bohan Li, Hao Zhao, Chaojian Li

arXiv 2608.17420首次发表:更新:

发表机构

Institute for AI Industry Research (AIR), Tsinghua University; Zhejiang University; University of Wisconsin–Madison; Great Wall Motor Company Limited; Shanghai Jiao Tong University; The Hong Kong University of Science and Technology(清华大学人工智能产业研究院; 浙江大学; 威斯康星大学麦迪逊分校; 长城汽车股份有限公司; 上海交通大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SPVC框架,通过结构化、全景、视频及跨数据集修复原则,训练两阶段可控视频扩散模型,解决驾驶场景渲染的伪影问题,实现跨数据集的高效修复。

AI 中文摘要

驾驶场景重建与渲染,尤其是基于3D高斯溅射(3D Gaussian Splatting)的技术,已成为自动驾驶仿真的重要组成部分。然而,在 ego 轨迹外推及场景编辑操作下,渲染视图常出现结构模糊、时间闪烁、前景-背景错位等问题。现有修复方法通常针对特定场景设计,如图像级新视角修复或物体编辑修正。本文提出SPVC,一种面向跨数据集驾驶场景渲染的结构化全景视频修复框架,其名称涵盖四项设计原则:1. 结构化修复:使用显式空间条件,包括相机位姿、3D边界框、高精地图(HD maps)引导修复过程,减少无控制的幻觉生成;2. 全景修复:同时修正背景渲染伪影(如道路、建筑、车道扭曲)与场景编辑引入的前景车辆伪影(如物体外观不一致);3. 视频修复:模型处理驾驶序列而非孤立帧,使伪影修正时可利用时间线索;4. 跨数据集修复:单个共享网络在多个驾驶数据集上训练并应用,减少对特定数据集或场景修复器的需求。具体而言,本文通过模拟欠约束的3DGS渲染及前景车辆插入伪影构建成对的退化-干净训练数据,训练了两阶段可控视频扩散模型,先处理视频级外观,再用结构化控制优化场景布局。

英文摘要

Driving scene reconstruction and rendering, especially with 3D Gaussian Splatting, has become an important component of autonomous driving simulation. However, rendered views often degrade under extrapolated ego trajectories and scene edits, producing blurry structures, temporal flicker, and foreground-background misalignment. Existing refinement methods are commonly designed for a specific setting, such as image-level novel-view repair or object-editing correction. In this paper, we introduce SPVC, a structured and panoptic video fixing framework for cross-dataset driving scene rendering. The name summarizes four design principles. (1) Structured fixing denotes the use of explicit spatial conditions, including camera pose, 3D bounding boxes, and HD maps, to guide the repair process and reduce uncontrolled hallucination. (2) Panoptic fixing refers to correcting both background rendering artifacts, such as distorted roads, buildings, and lanes, and foreground vehicle artifacts introduced by scene editing, such as inconsistent object appearance. (3) Video fixing means that the model operates on driving sequences rather than isolated frames, allowing temporal cues to be used during artifact correction. (4) Cross-dataset fixing means that a single shared network is trained and applied across multiple driving datasets, reducing the need for dataset-specific or scene-specific fixers. Concretely, we construct paired degraded-clean training data by simulating under-constrained 3DGS rendering and foreground vehicle insertion artifacts, and train a two-stage controllable video diffusion model that first addresses video-level appearance and then refines scene layout with structured controls.

CommentsProject page: https://li00147.github.io/SPVC-Project-Page/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑