AI 中文总结
针对现有道路重建方法需逐场景训练及场景依赖覆盖设计的局限,提出RoadVGGT框架,利用几何基础模型结合多视图图像等,经高斯头预测、坐标对齐及融合等重建紧凑高斯路面,提升多方面性能,证明几何基础模型在路面重建中的潜力。
AI 中文摘要
大规模路面重建支持高清地图绘制、自动驾驶感知、标注及模拟。现有道路专用优化方法虽能生成高质量道路表示,但需逐场景训练且依赖驾驶轨迹周围的场景覆盖设计,限制了新采集道路的可扩展重建。为此,我们引入了RoadVGGT,这是一个基于道路结构感知的前馈框架,无需测试时的逐场景优化即可重建紧凑的高斯路面。RoadVGGT利用几何基础模型结合提供的姿态和深度观测来利用多视图图像,并通过学习的高斯头预测密集的像素对齐高斯属性。为使这些密集预测可用于大型路面,我们将它们对齐到一致的度量世界坐标系中,并通过置信加权网格融合在道路对齐的XY平面上融合冗余高斯。类别感知分组和路侧人行道交界处保护进一步控制了易受影响道路结构周围的融合。由此产生的表示支持RGB和语义鸟瞰图、高程估计和新视图合成。RoadVGGT消除了现有方法中逐场景优化的需求,用紧凑的高斯表示重建完整路面,并提高了图像质量、语义映射和高程精度。大量实验证明了几何基础模型在可扩展前馈路面重建中的潜力。
英文摘要
Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-specialized optimization methods can produce high-quality road representations, but they typically require per-scene training and scene-dependent coverage design around the driving trajectory, limiting scalable reconstruction over newly collected roads. To address these limitations, we introduce RoadVGGT, a road-structure-aware feed-forward framework that reconstructs compact Gaussian road surfaces without test-time per-scene optimization. RoadVGGT uses a geometric foundation model to exploit multi-view images together with provided pose and depth observations, and predicts dense pixel-aligned Gaussian attributes through a learned Gaussian head. To make these dense predictions usable for large road surfaces, we align them into a consistent metric world coordinate system and fuse redundant Gaussians on the road-aligned XY plane through confidence-weighted grid fusion. Category-aware grouping and road--sidewalk junction protection further control fusion around vulnerable road structures. The resulting representation supports RGB and semantic bird's-eye-view maps, elevation estimation, and novel view synthesis. RoadVGGT eliminates the need for per-scene optimization in prior methods, reconstructs complete road surfaces with a compact Gaussian representation, and improves image quality, semantic mapping, and elevation accuracy. Extensive experiments demonstrate the potential of geometric foundation models for scalable feed-forward road surface reconstruction.