Geo-Embed:面向城市理解的统一多模态嵌入
Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding
浏览论文内容
中文总结 AI 辅助
本文针对现有多模态嵌入模型难以支持异构地理空间任务的问题,提出GeoMEB基准与Geo-Embed模型,在GeoMEB上实现15.3%的相对性能提升,为地理空间嵌入器发展提供方向。
中文摘要 AI 辅助
地理空间与城市应用日益需要模型对比街景图像、遥感观测、文本描述、区域提议、时间变化线索等异构证据,但现有多模态嵌入模型与基准大多围绕通用图像-文本匹配设计和评估,尚不明确统一嵌入空间能否支持涉及空间关系、细粒度语义、时间变化的异构地理空间任务。为填补这一空白,本文做出三项关键贡献:第一,提出GeoMEB——大规模多模态嵌入基准,标准化了检索、视觉问答、变化检测、分类、视觉定位共45项城市评估任务,同时包含132万条示例的训练集与28.6万条评估查询。第二,提出Geo-Embed——统一嵌入模型,将共享视觉-语言主干适配为指令条件下的异构地理空间输入(包括单图像、多图像、文本、区域、掩码)的查询-目标匹配。在GeoMEB上,Geo-Embed在代表性多模态嵌入器中取得最强综合性能,较最强基线实现15.3%的相对提升。这些结果为未来围绕显式查询-目标关系(包括语义、跨视图、区域级、时间对应)组织训练与评估的地理空间嵌入器提供了动力。
英文摘要
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes. To address this gap, we make three key contributions. First, we introduce GeoMEB, a large-scale multimodal embedding benchmark that standardizes 45 urban evaluation tasks across retrieval, visual question answering, change detection, classification, and visual grounding, together with training collections comprising 1.32M examples and 286K evaluation queries. Second, we present Geo-Embed, a unified embedding model that adapts a shared vision-language backbone to instruction-conditioned query-target matching over heterogeneous geospatial inputs, including single images, multiple images, text, regions, and masks. On GeoMEB, Geo-Embed achieves the strongest overall performance among representative multimodal embedders, with a 15.3% relative improvement over the strongest baseline. These results motivate future geospatial embedders that organize training and evaluation around explicit query-target relations, including semantic, cross-view, region-level, and temporal correspondence.