arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一次查询,多尺度:用于高效分层跨视图地理定位的稀疏专家混合模型

One Query, Many Scales: Sparse Mixture-of-Experts for Efficient Hierarchical Cross-View Geo-Localization

Ruijie Fan, Junyan Ye, Qi Zhu, Weijia Li

arXiv 2608.01060首次发表:更新:

AI 中文总结

GeoMoE作为稀疏专家混合双编码器,将全局多尺度表示学习与局部分层搜索解耦,在跨视图地理定位任务中提升了精度、效率与跨分辨率迁移性能。

AI 中文摘要

跨视图地理定位(CVGL)是指为地面视图查询检索带有地理标签的卫星图像。大多数系统会对平面、固定分辨率的图库进行穷举搜索,这在大区域范围内成本很高,且难以适应卫星分辨率的变化。自回归的由粗到精替代方案减少了比较次数,但会将后续预测与早期决策以及预定义的层级绑定。我们提出GeoMoE,一种稀疏专家混合双编码器,它将全局多尺度表示学习与局部分层搜索解耦。全局多尺度监督和内容自适应路由将地面和卫星图像跨分辨率映射到全局可比的嵌入空间。推理时,每张图像仅编码一次,概率波束搜索会遵循父子链接对小型候选子集进行评分。后续层级复用这些描述符,而非前序层级生成的特征,从而限制了特征级误差传播和层级耦合。我们还提出VIGOR-M,一个包含四个城市的基准,具有显式的父子卫星层级,并保留了半步骤图库,用于单分辨率、跨分辨率和分层评估。GeoMoE在Just Zoom In上达到95.78%的R@40m,比之前的最佳结果高出2.77个百分点;在VIGOR-M上达到62.39%的R@1,其描述符匹配的计算量为0.885 MMAC/查询,仅为穷举L3扫描的5.27%,同时在R@1上比最强的穷举基线高出3.12个百分点。在L1、L2和L3上训练的单个模型,在所有六个图库上的表现均优于匹配的密集对照模型,并能迁移到三个保留的分辨率。通过将全局训练的嵌入与局部分层搜索解耦,GeoMoE同时提升了定位精度、搜索效率和跨分辨率迁移能力。

英文摘要

Cross-view geo-localization (CVGL) retrieves geo-tagged satellite imagery for a ground-view query. Most systems exhaustively search a flat, fixed-resolution gallery, incurring high cost over large areas and adapting poorly to satellite resolution changes. Autoregressive coarse-to-fine alternatives reduce comparisons but bind later predictions to earlier decisions and a predefined hierarchy. We introduce GeoMoE, a sparse mixture-of-experts dual encoder that decouples global multi-scale representation learning from local hierarchical search. Global multi-scale supervision and content-adaptive routing map ground and satellite images across resolutions into a globally comparable embedding space. At inference, each image is encoded once, and probabilistic beam search follows parent--child links to score a small candidate subset. Later levels reuse these descriptors rather than features generated by preceding levels, limiting feature-level error propagation and hierarchy coupling. We further introduce VIGOR-M, a four-city benchmark with an explicit parent--child satellite hierarchy and held-out half-step galleries for single-resolution, cross-resolution, and hierarchical evaluation. GeoMoE achieves 95.78% R@40m on Just Zoom In, 2.77 percentage points above the previous best, and 62.39% R@1 on VIGOR-M. The latter requires 0.885 MMAC/query for descriptor matching, 5.27% of an exhaustive L3 scan, while exceeding the strongest exhaustive baseline by 3.12 percentage points in R@1. One model trained on L1, L2, and L3 also outperforms a matched dense control across all six galleries and transfers to three withheld resolutions. By decoupling globally trained embeddings from local hierarchical search, GeoMoE jointly improves localization accuracy, search efficiency, and cross-resolution transfer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑