面向自动驾驶的基于几何的统一3D感知
Geometry-Grounded Unified 3D Perception for Autonomous Driving
浏览论文内容
中文总结 AI 辅助
该研究针对现有自动驾驶3D感知框架无法保留显式度量几何的问题,提出GeoUP框架,通过分解跨图像交互并注入射线图编码实现统一3D感知,在多数据集上达到SOTA性能。
中文摘要 AI 辅助
基于相机的自动驾驶感知需要一种共享表示,该表示能在同步多相机流中保留度量3D结构。然而,现有的基于图像的框架通常依赖于为语义识别预训练的骨干网络,并通过下游特定任务模块引入3D几何,导致其共享表示可能无法保留显式度量几何和一致的3D场景结构。本文提出了一种基于几何的统一3D感知(Geometry-grounded Unified 3D Perception, GeoUP)框架,该框架将面向重建的VGGT潜变量适配到已校准的流式多相机驾驶场景中。GeoUP将跨图像交互分解为自注意力、时间注意力和视图注意力,以捕获结构上不同的时间和跨视图对应关系;进一步注入感知校准的射线图编码,以提供度量尺度和相机几何。生成的基于几何的潜变量被解码为度量深度估计、3D目标检测和语义占据预测,对应于同一3D场景的表面级、实例级和体积级输出。通过联合多任务和多数据集训练,GeoUP有效利用异构标注,并在不同传感器配置和感知范围内实现泛化。在nuScenes、Argoverse 2、Waymo、KITTI和DDAD上的大量实验表明,GeoUP在检测、占据和深度估计方面达到了SOTA性能,这些结果验证了基于几何的表示在统一3D驾驶感知中的有效性。
英文摘要
Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pretrained for semantic recognition, and introduce 3D geometry through downstream task-specific modules. As a result, their shared representations may fail to preserve explicit metric geometry and consistent 3D scene structure. In this paper, we present a Geometry-grounded Unified 3D Perception (GeoUP) framework that adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes. GeoUP factorizes cross-image interaction into self, temporal, and view attention to capture structurally distinct temporal and cross-view correspondences. It further injects calibration-aware raymap encodings to provide metric scale and camera geometry. The resulting geometry-grounded latent is decoded for metric depth estimation, 3D object detection, and semantic occupancy prediction, corresponding to surface-, instance-, and volume-level readouts of the same 3D scene. Through joint multi-task and multi-dataset training, GeoUP effectively leverages heterogeneous annotations and generalizes across diverse sensor configurations and perception ranges. Extensive experiments on nuScenes, Argoverse 2, Waymo, KITTI, and DDAD demonstrate that GeoUP achieves SOTA performance across detection, occupancy, and depth estimation. These results validate the effectiveness of geometry-grounded representations for unified 3D driving perception.