发表机构
University College London; Karlsruhe Institute of Technology; Hunan University; University of Alberta; Shenzhen University(伦敦大学学院; 卡尔斯鲁厄理工学院; 湖南大学; 阿尔伯塔大学; 深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对跨视角视频地理定位的局限,提出X²Localizer跨粒度对齐框架与SWRL策略,实现渐进式跨视角视频地理定位,大幅提升早期定位性能,缩小基准评估与实际部署的差距。
AI 中文摘要
跨视角视频地理定位(Cross-view Video Geo-localization,CVG)旨在通过检索对应带地理标签的航拍图像来定位地面视角视频。然而,CVG方法依赖固定长度输入和事后优化,阻碍了在部分或动态观测下面向在线的定位。在本研究中,我们将渐进式跨视角视频地理定位(Progressive Cross-view Video Geo-localization,PCVG)表述为CVG的面向部署的扩展与评估协议,支持在不同时间预算、基于前缀的推理、随机起始评估以及带中断的长距离定位下的定位。为探索PCVG,我们引入X²Localizer,这是一种跨粒度对齐框架,以依赖预算的非对称目标共同监督全局前缀到航拍图像的检索,以及标记聚合的帧-航拍图像-瓦片匹配。此外,我们引入滑动窗口重定位(Sliding-Window Re-Localization,SWRL)策略,该策略动态刷新候选区域以用于故障恢复和长距离部署,无需重新处理完整序列。大量实验表明,X²Localizer保留了传统全视频的性能,仅带来+0.1的Recall@1和+0.3的Recall@10的微小提升,同时大幅改善早期定位。在具有挑战性的单帧设置中,X²Localizer比之前的最先进方法将粗粒度检索的Recall@1提升了+4.7,Recall@10提升了+11.5。结合SWRL,我们的方法进一步在随机起始和长距离场景下实现了稳健的渐进式定位,缩小了基准评估与实际部署之间的差距。
英文摘要
Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on fixed-length inputs and post-hoc refinement, hindering online-oriented localization under partial or dynamic observations. In this work, we formulate Progressive Cross-view Video Geo-localization (PCVG) as a deployment-oriented extension and evaluation protocol of CVG, enabling localization under varying temporal budgets, prefix-based inference, random-start evaluation, and long-range localization with interruptions. To explore PCVG, we introduce X$^2$Localizer, a cross-grained alignment framework that jointly supervises global prefix-to-aerial retrieval and token-aggregated frame--aerial-tile matching with a budget-dependent asymmetric objective. Furthermore, we introduce a Sliding-Window Re-Localization (SWRL) strategy that dynamically refreshes candidate regions for failure recovery and long-range deployment without full-sequence reprocessing. Extensive experiments show that X$^2$Localizer preserves conventional full-video performance, with marginal gains of +0.1 Recall@1 and +0.3 Recall@10, while substantially improving early localization. In the challenging single-frame setting, X$^2$Localizer improves coarse retrieval by +4.7 Recall@1 and +11.5 Recall@10 over the previous state-of-the-art method. With SWRL, our approach further enables robust progressive localization under random-start and long-distance scenarios, narrowing the gap between benchmark evaluation and real-world deployment.
CommentsAccepted to The 37th British Machine Vision Conference (BMVC 2026)