CrossDepth:用于通用多视角环绕深度估计的几何约束注意力机制
CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation
浏览论文内容
中文总结 AI 辅助
本文针对多视角环绕深度估计的跨图像不一致问题,提出几何约束注意力机制的CrossDepth模型,经DDAD、nuScenes数据集评估,其深度精度与一致性优于现有最优自监督方法。
中文摘要 AI 辅助
对周围环境的可靠3D理解是自动驾驶的核心需求。多视角环绕相机 rig 能提供广泛的场景覆盖,但空间相邻的图像通常仅存在极少重叠。因此,大多数像素的深度必须从单目外观线索推断,而这些线索在不同图像中可能呈现不同表现,导致深度估计模型对其产生不同解读。本文针对跨图像不一致性的两个主要来源:相机内参差异及各图像有限的感受野,提出对应解决方案:前者通过将特征基于逐像素的相机感知射线嵌入进行条件化,使网络能考虑单目线索中依赖相机的变化;后者通过受几何合理区域约束的跨图像注意力,将每个像素的上下文扩展至其自身图像之外,该几何合理区域由校准后的相机 rig 装置推导得出。该模型基于光度一致性以完全自监督方式训练,在 DDAD 和 nuScenes 数据集上的评估显示,在域内和跨域评估场景下,其相比现有最优自监督方法,整体深度精度和跨图像深度一致性均有所提升。代码可在提供的 URL 获取。
英文摘要
Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-view surround camera rigs provide broad scene coverage, but the spatially adjacent images typically overlap only minimally. Consequently, the depth of most pixels must be inferred from monocular appearance cues. These cues can appear differently across images and may therefore be interpreted differently by the depth estimation model. We target two main sources of cross-image inconsistency: differences in camera intrinsics and the limited receptive field of each image. We address the former by conditioning the features on per-pixel camera-aware ray embeddings, enabling the network to account for camera-dependent variations in monocular cues. We address the latter by extending each pixel's context beyond its own image through cross-image attention constrained to geometrically plausible regions, derived from the calibrated rig setup. The model is trained in a fully self-supervised manner based on photometric consistency. Evaluations on DDAD and nuScenes show improved overall depth accuracy and cross-image depth consistency over state-of-the-art self-supervised methods under in-domain and cross-domain evaluation. Code is available at https://abualhanud.github.io/CrossDepthPage/.
发表机构
- Leibniz University Hannover(汉诺威莱布尼茨大学)
机构由 AI 辅助整理,请以论文原文为准。