发表机构
University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究无人机跨域全局视觉定位问题,提出检索匹配框架RIM,通过采样参考视图、两阶段微调SALAD及冻结检索器蒸馏解码器等方法,在多数据集零样本评估中表现出色,建立了高效可部署的定位管道。
AI 中文摘要
利用遥感参考地图进行无人机全局视觉定位备受关注。但无人机与参考图像在采集时间和成像平台上的差异,导致跨域外观和视角变化,给六自由度姿态估计带来挑战。我们通过从谷歌3D瓦片跨位置、高度和方向采样无人机视角参考视图来解决这些变化。采用两阶段跨域微调方法,利用姿态相近的正样本和地理上遥远的硬负样本对SALAD进行调整,同时通过局部几何一致性对Top-K候选进行重新排序。还提出检索匹配(RIM),冻结适配的DINOv2-B检索器,蒸馏出局部描述符解码器。在重建的EPFL Urbanscape和自采的长安公园数据集上进行零样本评估,RIM优于十个近期检索基线系列。在全3D距离度量下,在25/50米处,RIM在EPFL上比SALAD的Recall@1提高8.55/13.77个百分点,在公园数据集上提高4.45/8.94个百分点。在Top-K = 5时,完整的定位查询端到端耗时67.9毫秒,比最强的单独稀疏匹配基线快1.8倍,比RoMa快40多倍,同时实现可比的重新排序精度。这些结果为GNSS受挑战环境下的无人机全局视觉定位建立了高效且可部署的管道。
英文摘要
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, differences in acquisition time and imaging platform between UAV and reference imagery introduce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation. To mitigate these shifts, we render UAV-viewpoint references from Google 3D Tiles across locations, altitudes, and orientations. A two-stage strategy adapts SALAD with pose-near positives and geographically distant hard negatives; local geometric consistency then re-ranks the Top-K candidates. We further propose Retrieval-In-Matching (RIM), which freezes the adapted DINOv2-B retriever and distills a local-descriptor decoder from its token field and a shallow VGG19 detail stream. One query-side DINOv2-B backbone forward therefore supports both SALAD retrieval and local description, eliminating a second foundation-model backbone while preserving the retrieval descriptors by construction. We evaluate RIM zero-shot on the reconstructed EPFL Urbanscape and self-collected Chang'an Park datasets, both geographically disjoint from the training data. RIM outperforms ten retrieval baselines. Under the full 3D distance metric at 25/50 m, it improves Recall@1 over SALAD by 8.55/13.77 percentage points on EPFL and 4.45/8.94 points on Park. At Top-K=5, the measured online query path through retrieval, candidate matching, and robust geometric verification takes 90.8 ms: 1.2 times faster than the strongest separate sparse-matching baseline and over 30 times faster than RoMa, while maintaining comparable re-ranking accuracy. These results demonstrate an efficient UAV global visual localization pipeline under unreliable satellite navigation. The source code is available at https://github.com/curious-energy/RIM.
CommentsAccepted for publication in Knowledge-Based Systems. 62 pages, including supplementary material
Journal refKnowledge-Based Systems 352 (2026) 117027
DOI:10.1016/j.knosys.2026.117027