arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于昼夜无人机视角地理定位的统一基准和模态自适应网络

A Unified Benchmark and Modality-Adaptive Network for Day-and-Night Drone-View Geo-Localization

Songtianhao Xu, Zhongwei Chen, Zhao-Xu Yang, Weifeng Wang

arXiv 2607.25778首次发表:更新:

发表机构

Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有无人机视角地理定位基准局限性,构建IRCHN基准,提出MASTR - Net框架,能统一评估不同光照条件及传感模态下的DVGL,该框架在相关基准测试中表现出色。

AI 中文摘要

大多数现有的无人机视角地理定位(DVGL)基准包含在单一光照条件下捕获的无人机图像,且缺乏来自相同位置的地理对齐的可见无人机图像、红外无人机图像和卫星图像。为评估DVGL方法在具有挑战性光照条件下的泛化能力,一些方法在可见基准上训练模型并在独立红外基准上测试,这难以在统一基准中系统评估DVGL。为此构建了IRCHN基准,它包含来自四个场景类别的26460张图像,能统一评估DVGL方法。还提出了MASTR - Net框架,集成了模态自适应特征增强等技术。实验表明MASTR - Net在IRCHN上优于现有方法,并在其他红外基准上有竞争力。

英文摘要

Most existing drone-view geo-localization (DVGL) benchmarks contain drone imagery captured under a single illumination condition and lack geographically aligned visible drone images, infrared drone images, and satellite images from the same locations. To evaluate the generalization capability of DVGL methods under challenging illumination conditions, some methods train models on a visible benchmark and test them on an independent infrared benchmark. This protocol essentially constitutes transfer between datasets, which makes it difficult to systematically evaluate DVGL across daytime and nighttime conditions within a unified benchmark. To address this limitation, we construct IRCHN,a real-world DVGL benchmark designed for localization across different illumination conditions. IRCHN contains 26,460 images collected from 8,820 geographic locations across four representative scene categories, including farmland, coastline, forest, and urban areas. Each location provides one visible drone image, one infrared drone image, and one corresponding satellite image, which enables unified evaluation of DVGL methods across different illumination conditions and sensing modalities. We further propose the Modality-Adaptive State-Space Transport Relation Network (MASTR-Net), a DVGL framework tailored to localization under varying illumination conditions. MASTR-Net integrates modality-adaptive feature enhancement, bidirectional selective state-space relation modeling, and soft optimal transport relation alignment to jointly reduce modality gaps and view-induced structural discrepancies. Extensive experiments demonstrate that MASTR-Net outperforms existing state-of-the-art methods on IRCHN for localization under varying illumination conditions and achieves competitive performance on two infrared benchmarks, IR-VL328 and CVGL-RGBT. Code: https://github.com/SongtianhaoXu/MASTR-Net

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑