发表机构
The Hong Kong Polytechnic University; Southern University of Science and Technology; Wuhan University; The Chinese University of Hong Kong; Huawei Technologies Co., Ltd(香港理工大学; 南方科技大学; 武汉大学; 香港中文大学; 华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对卫星-地面定位中的几何与描述符歧义,提出GeoSem-BEV方法,通过径向深度、垂直高度和共享语义监督约束BEV表示学习,显著降低方向误差。
AI 中文摘要
卫星-地面定位旨在估计地面相机在地理参考卫星图像中的平面位置和偏航方向。最近的方法将地面和卫星特征映射到共享的鸟瞰图(BEV)空间中,并建立空间对应关系。然而,深度约束不足可能将一个地面特征分配到沿观察方向的不同距离处,从而在BEV特征放置中产生几何歧义。不同位置处的相似外观也可能导致描述符匹配歧义,而现有的描述符学习缺乏显式的语义监督来区分它们。我们提出了GeoSem-BEV,一种几何-语义约束的BEV表示学习方法。径向深度监督约束距离分配,垂直高度监督约束高度聚合。共享的显式语义监督促进跨视图的语义预测一致性,并有助于区分具有相似语义的位置。这些约束改善了特征放置和描述符可区分性,增强了最先进的BEV定位模型。在具有未知方向的VIGOR上,GeoSem-BEV在跨区域和同区域设置中分别将平均方向误差降低了37.2%和38.1%。在DReSS-D上,相应的误差分别降低了10.8%和15.6%。在KITTI-CVL上,在10度方向噪声下,它将同区域平均方向误差降低了26.8%。
英文摘要
Satellite-ground localization estimates the planar position and yaw orientation of a ground camera within a geo-referenced satellite image. Most recent methods map ground and satellite features into a shared bird's-eye-view (BEV) space and establish spatial correspondences. However, insufficient depth constraints can assign one ground feature to different distances along a viewing direction, creating geometric ambiguity in BEV feature placement. Similar appearances at different locations can also create descriptor matching ambiguity, while existing descriptor learning lacks explicit semantic supervision to distinguish them. We propose GeoSem-BEV, a geometry-semantic constrained BEV representation learning method. Radial depth supervision constrains distance assignment, and vertical height supervision constrains height aggregation. Shared explicit semantic supervision promotes consistent semantic predictions across views and helps distinguish locations with similar semantics. These constraints improve feature placement and descriptor discriminability, enhancing state-of-the-art BEV localization models. On VIGOR with unknown orientation, GeoSem-BEV reduces mean orientation error by 37.2% and 38.1% in the cross-area and same-area settings, respectively. The corresponding errors are reduced by 10.8% and 15.6% on DReSS-D. On KITTI-CVL, it reduces same-area mean orientation error by 26.8% under 10 degree orientation noise.
Comments10 pages, 2 figures, and 4 tables