arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ARC-Loc:利用方位角射线收敛性作为直接跨视角定位的几何线索

ARC-Loc: Leveraging Azimuthal Ray Convergence as a Geometric Cue for Direct Cross-View Localization

Hyeongsik Kim, Mincheol Kim, Heejoon Moon, Je Hyeong Hong

arXiv 2609.04965首次发表:更新:

发表机构

Hanyang University(汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出ARC-Loc方法,无需依赖BEV变换和外部深度基础模型,利用方位角射线收敛的几何约束实现跨视角定位,在VIGOR和KITTI数据集上保持了有竞争力的精度,且推理更快、内存效率更高。

AI 中文摘要

跨视角定位(CVL)通过将地面图像与地理参考的卫星图像匹配来估计地面图像的位姿。为弥合极端视角差距,主流流程依赖鸟瞰图(BEV)变换或2D转3D提升,但从单张地面图像推导3D结构本质上是不适定的,导致这些方法在3D提升或BEV投影过程中承受几何失真与计算成本;此外,依赖外部深度基础模型解决该问题会引入延迟,且仍易受噪声预测影响。本研究提出一种受人类导航技术“后方交会”启发的不同方法,可在不依赖外部深度基础模型的情况下直接完成地面到卫星图像的匹配与定位。该方法的关键见解为:(i)地面关键点可转换为卫星地图上的方位角射线;(ii)这些射线理想情况下会在用户位置收敛。通过利用该几何约束进行直接线点对应,我们引入了最小化的方位角射线收敛(ARC)求解器来识别交点,同时提出ARC损失以优化匹配网络。通过消除对计算密集型BEV变换和外部深度基础模型的依赖,本方法实现了更快、内存高效的推理,且其显式特征匹配确保了与现有框架的直接兼容性。在VIGOR和KITTI上的实验表明,ARC-Loc与近期方法相比保持了有竞争力的定位精度,凸显了其实用性。

英文摘要

Cross-view localization (CVL) estimates the pose of a ground image by matching it to a geo-referenced satellite image. To bridge the extreme viewpoint gap, mainstream pipelines rely on Bird's-Eye-View (BEV) transformations or 2D-to-3D lifting. However, deriving 3D structures from a single ground image is fundamentally ill-posed, causing these methods to endure geometric distortions and computational costs during 3D lifting or BEV projection. Furthermore, relying on external depth foundation models to resolve this introduces latency and remains susceptible to noisy predictions. In this work, we present a different approach inspired by a human navigation technique called resection, that can perform direct ground to satellite image matching and localization without relying on external depth foundation models. The key insights of our method are that (i) ground keypoints can be translated into azimuthal rays on the satellite map, and (ii) these rays ideally converge at the user location. Exploiting this geometric constraint through direct line-to-point correspondences, we introduce a minimal Azimuthal Ray Convergence (ARC) solver to identify the intersection, alongside an ARC loss to optimize the matching network. By eliminating dependencies on computationally heavy BEV transformations and external depth foundation models, our approach achieves faster, memory-efficient inference, while its explicit feature matching ensures straightforward compatibility with existing frameworks. Experiments on VIGOR and KITTI demonstrate that ARC-Loc maintains competitive localization accuracy compared to recent approaches, highlighting its practicality.

CommentsAccepted to ECCV2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑