AI 中文总结
研究仅视觉3D驾驶占用预测问题,提出VGOcc方法,通过从基础模型学习视觉和几何线索,纳入高斯建模的初始化与细化,实现高效语义占用预测,在nuScenes实验中达最优性能。
AI 中文摘要
仅视觉占用预测需要从校准的环视图像中恢复语义3D占用场,每个视图沿相机光线提供深度模糊的观测。现有方法从密集结构表示发展到稀疏高斯基元,提高了3D场景表示效率,但高斯学习仍主要依赖图像域特征,几何信息有限。本文提出VGOcc,从基础模型学习视觉和几何线索用于高斯建模,将其纳入基元初始化和细化,产生适用于语义占用预测的视觉几何高斯表示。具体包括视觉几何高斯诞生、姿态感知特征学习等步骤。在nuScenes上的实验表明VGOcc在仅视觉3D占用预测中达到了当前最优性能。
英文摘要
Vision-only occupancy prediction requires recovering a semantic 3D occupancy field from calibrated surround-view images, where each view provides observations with ambiguous depth along camera rays. Existing methods have progressed from dense structured representations to sparse Gaussian primitives, improving the efficiency of 3D scene representation. However, Gaussian learning still relies primarily on image domain features, which provide limited explicit geometric information for volumetric reasoning. Our key observation is that effective Gaussian occupancy modeling requires not only sparse primitives, but also richer geometric and semantic learning cues. In this paper, we propose VGOcc, which learns visual and geometric cues from foundation models for Gaussian modeling. VGOcc incorporates these cues into primitive initialization and refinement, yielding a representation termed Visual-Geometric Gaussians tailored to semantic occupancy prediction. Specifically, we propose Visual-Geometric Gaussian Birth to form spatially balanced Gaussian centers from ray depth hypotheses, while visual semantic features initialize primitive attributes. Next, we design Pose-Aware Feature Learning to combine foundation tokens with camera embeddings and calibrated ray information. Features from neighboring views are then aggregated at projected 3D locations for each Gaussian refinement stage. Finally, Gaussian decoder refines birth Gaussians with pose-aware features and renders them into semantic occupancy. Experiments on nuScenes demonstrate that VGOcc achieves state-of-the-art performance in vision-only 3D occupancy prediction. Codes will be available at https://github.com/JHLin42in/VGOcc.