arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向单帧环视驾驶场景重建的视觉几何基础感知高斯模型

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

Junhong Lin, Jinlong Wang, Xianda Guo, Yanlun Peng, Wei Zheng, Guoqing Liu, Hanli Wang, Tiesong Zhao, Wei Gao

arXiv 2608.10682首次发表:更新:

发表机构

Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology; School of Electronic and Computer Engineering, Peking University; Peng Cheng Laboratory; Wuhan University; Great Wall Motor; Minieye Corporation; Tongji University; Fuzhou University(广东省超高清沉浸式媒体技术重点实验室; 北京大学电子与计算机工程学院; 鹏城实验室; 武汉大学; 长城汽车; Minieye公司; 同济大学; 福州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出VGGD框架,通过引入视觉几何先验、双路径颈部等模块,在nuScenes基准上实现了单帧环视驾驶场景重建的最优渲染质量与几何一致性。

AI 中文摘要

单帧环视重建因相机间重叠度极低而面临严重的几何不稳定性与渲染伪影问题。现有方法虽依赖复杂解码器或辅助线索,但仍受限于上游特征的几何能力薄弱。本文认为,利用预训练视觉几何先验可增强上游表示,缓解稀疏环视视图中的几何歧义。为此,我们提出VGGD——一种面向前馈式环视驾驶重建的视觉几何基础感知3D Gaussian Splatting框架,其将几何建模转移至前端,并使基础先验适配驾驶相机设置。首先,VGGD利用VGGT提供可迁移的多视图几何先验令牌;其次,引入双路径颈部模块以解耦几何一致性与外观感知表示,提升弱观测区域的外观补全效果;进一步应用尺度预热策略,以稳定早期几何学习并抑制自车姿态变化下的尺度漂移;最终采用混合像素-体积高斯解码器生成可渲染的3D高斯场景,用于新视图合成。在nuScenes单帧基准上的实验表明,VGGD在对比方法中实现了最优的整体渲染质量,并提升了相对几何一致性。

英文摘要

Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing methods rely on complex decoders or auxiliary cues, they remain bottlenecked by the weak geometric capacity of upstream features. We argue that leveraging pretrained visual geometry priors strengthens upstream representations and alleviates the geometric ambiguity in sparse surround views. To this end, we propose VGGD, a visual geometry foundation-aware 3D Gaussian Splatting framework for feed-forward surround-view driving reconstruction, which shifts geometric modeling to the frontend and adapts foundation priors to the driving camera setting. First, VGGD leverages VGGT to provide transferable multi-view geometric prior tokens. Next, we introduce a Dual-Path Neck to decouple geometry-consistent and appearance-aware representations, improving appearance completion in weakly observed regions. We further apply Scale Warmup to stabilize early geometry learning and suppress scale drift under ego-pose changes. Finally, we use a hybrid pixel--volume Gaussian decoder to produce a renderable 3D Gaussian scene for novel-view synthesis. Experiments on the nuScenes single-frame benchmark show that VGGD achieves the best overall rendering quality among the compared methods and improves relative geometric consistency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑