USR-Drive:通过3D高斯与边界框联合去噪实现统一驾驶场景表示
USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes
浏览论文内容
中文总结 AI 辅助
本文提出USR-Drive框架,将3D高斯与边界框作为对齐的潜在令牌流,通过统一多模态扩散Transformer联合去噪,在nuScenes和VKitti数据集上实现动态重建与3D检测的SOTA性能。
中文摘要 AI 辅助
自动驾驶的空间表示学习旨在将原始视觉信号映射为结构化的3D场景表示,其中以对象为中心的边界框和面向渲染的3D基元(如3D高斯)是场景理解的两个截然不同但高度互补的层级。现有方法通常将动态重建和实例级感知视为独立任务,尽管二者的共同目标是估计底层3D世界状态,这导致动态重建的约束不足,而3D检测缺乏几何基础。为解决这一差距,本文提出USR-Drive,这是一种统一的条件生成框架,仅给定经过位姿估计的多视图驾驶视频,即可在共享场景表示中联合恢复密集动态几何和实例级对象布局。具体而言,USR-Drive将密集高斯基元和稀疏3D边界框表示为两个对齐的潜在令牌流,并通过统一的多模态扩散Transformer对其进行联合去噪。与先前将边界框作为外部条件或通过独立模块预测的范式不同,USR-Drive将它们视为相互约束的状态变量,并采用统一位置编码(UPE)在共享度量时空坐标中对齐异构令牌。通过这种统一的表示和生成框架,两种模态相互增强:几何为边界框预测提供密集度量证据,而边界框提供实例级结构先验,有助于保持空间一致性并减少序列3D几何表示中的歧义。本文方法在nuScenes和VKitti数据集上的动态重建和3D检测任务均取得了最先进的结果。
英文摘要
Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians) serve as two distinct yet highly complementary levels for scene understanding. Existing methods typically treat dynamic reconstruction and instance-level perception as separate tasks, despite their shared goal of estimating the underlying 3D world state. As a result, dynamic reconstruction is under-constrained while 3D detection lacks geometric grounding. To address this gap, we propose USR-Drive, a unified conditional generative framework that, given only posed multi-view driving videos, jointly recovers dense dynamic geometry and instance-level object layouts within a shared scene representation. Specifically, USR-Drive represents dense Gaussian primitives and sparse 3D bounding boxes as two aligned latent token streams and jointly denoises them with a unified multi-modal diffusion Transformer. Unlike prior paradigms that use boxes as external conditions or predict them with detached modules, USR-Drive treats them as mutually constrained state variables with a Unified Positional Encoding (UPE) that aligns heterogeneous tokens within a shared metric spatiotemporal coordinate. Via such unified representation and generative framework, the two modalities reinforce each other: geometry supplies dense metric evidence for box prediction, while boxes provide instance-level structural priors that help preserve spatial consistency and reduce ambiguity in sequential 3D geometric representation. Our approach successfully delivers state-of-the-art results for both dynamic reconstruction and 3D detection on the nuScenes and VKitti datasets.