多视图专家混合与视觉-语言重排序用于跨视角目标地理定位
Multi-View Mixture-of-Experts with Vision-Language Reranking for Cross-View Object Geo-Localization
浏览论文内容
中文总结 AI 辅助
提出MVLGeo框架,通过视觉-语言重排序和多视图专家混合架构,统一多视角目标地理定位,减少参数冗余,提升跨视角知识共享,在基准上达到最先进性能。
中文摘要 AI 辅助
跨视角目标地理定位(CVOGL)利用无人机或街景查询在卫星图像中定位目标。现有方法为每个视角训练独立的检测器,导致参数冗余并阻碍跨视角知识共享。此外,排名靠前的卫星候选图像往往在视觉上相似,因此仅凭视觉外观和类别标签不足以解决这种歧义。为解决这些问题,我们提出了MVLGeo,一个旨在统一多视角并减少模型冗余的高效框架。首先,我们引入查询视角的环境上下文文本作为线索,通过视觉-语言重排序(VL-Rerank)来区分视觉上相似的候选。其次,我们设计了一种多视图专家混合架构(MV-MoE),采用共享编码器和视角特定专家以减少冗余并促进知识共享,同时通过跨视角对比学习对齐其表示以保持一致性。第三,我们引入自适应椭圆先验(ESAM-Prior)作为辅助位置编码,用于各向异性几何感知。在CVOGL基准上的大量实验证实,MVLGeo作为统一的多查询视角模型,达到了最先进的性能,展现出对输入退化的鲁棒性和跨视角的泛化能力。代码和模型将在GitHub上提供,以促进未来的工作。
英文摘要
Cross-view object geo-localization (CVOGL) locates a target in satellite imagery using drone or street-view queries. Existing methods train separate detectors for each viewpoint, leading to parameter redundancy and impeding cross-view knowledge sharing. Moreover, top-ranked satellite candidates are often visually similar, so visual appearance and categorical labels alone are insufficient to resolve such ambiguity. To address these, we propose MVLGeo, an efficient framework designed to unify multiple viewpoints and reduce model redundancy. First, we introduce environmental contextual text from the query view as cues to distinguish visually similar candidates via Vision-Language Reranking (VL-Rerank). Second, we design a multi-view Mixture-of-Experts architecture (MV-MoE) with a shared encoder and view-specific experts to reduce redundancy and promote knowledge sharing, while cross-view contrastive learning aligns their representations for consistency. Third, we introduce an adaptive elliptical prior (ESAM-Prior) as auxiliary positional encoding for anisotropic geometric perception. Extensive experiments on the CVOGL benchmarks confirm that MVLGeo, as a unified model for multiple query viewpoints, achieves state-of-the-art performance, demonstrating robustness to input degradation and generalization across viewpoints. Code and models will be available on GitHub to facilitate future work.
发表机构
- College of Computer Science, Beijing University of Technology(北京工业大学计算机科学学院)
- Zhengzhou University(郑州大学)
- Zhejiang University(浙江大学)
- Beijing Institute of Technology(北京理工大学)
- Beijing Electronic Science and Technology Institute(北京电子科技学院)
- Intelligent Science & Technology Academy of CASIC(中国航天科工集团智能科学技术研究院)
机构由 AI 辅助整理,请以论文原文为准。