AI 中文总结
针对现有语义场景补全方法的深度误差传播问题,提出RayLift框架,通过互补上下文编码器、深度射线证据提升模块和语义感知体素集成器,在两个基准数据集上实现优于现有方法的性能。
AI 中文摘要
基于相机的3D语义场景补全(Semantic Scene Completion, SSC)为自动驾驶和机器人技术提供全面的场景理解。然而,现有方法常将立体深度估计视为确定性几何约束,导致深度不确定性和局部对应误差直接传播到体素表示中。为解决该问题,我们提出RayLift框架,其以立体几何为度量参考,同时融入互补射线证据以自适应恢复可靠的3D结构。RayLift首先采用互补上下文编码器(Complementary Context Encoder),从冻结的3D视觉基础模型中提取几何感知先验,从而丰富场景上下文;随后引入深度射线证据提升模块(Depth Ray Evidence Lifter),联合建模几何差异、深度置信度和空间不确定性,以沿每条相机射线自适应采样并加权候选表面位置;最后,语义感知体素集成器(Semantic-Aware Voxel Integrator)通过显式建模空间支持,将所得射线证据注入体素特征中。在SemanticKITTI和SSCBench-KITTI-360上开展的大量实验表明,RayLift取得了具竞争力的性能,且始终优于现有方法。
英文摘要
Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing methods often treat stereo depth estimates as deterministic geometric constraints, causing depth uncertainty and local correspondence errors to propagate directly into voxel representations. To address this issue, we propose RayLift, a framework that uses stereo geometry as a metric reference while incorporating complementary ray evidence to recover reliable 3D structures adaptively. RayLift first employs a Complementary Context Encoder that extracts geometry-aware priors from a frozen 3D vision foundation model, thereby enriching the scene context. It then introduces a Depth Ray Evidence Lifter module that jointly models geometric dissimilarity, depth confidence, and spatial uncertainty to adaptively sample and weight candidate surface locations along each camera ray. Finally, a Semantic-Aware Voxel Integrator injects the resulting ray evidence into voxel features by explicitly modeling their spatial support. Extensive experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that RayLift achieves competitive performance and consistently outperforms existing methods.