发表机构
National Institute of Technology, Karnataka; Visa Inc; Greenlight(卡纳塔克邦国家理工学院; 维萨公司; Greenlight)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对地理空间基础模型现有预训练方法无法捕捉卫星场景组成性的问题,提出感知组成的预训练框架,在多任务上优于更大参数规模的现有模型,在 ForestNet-12 数据集上取得 55.6% 的相对 mAP@10 提升。
AI 中文摘要
地理空间基础模型已成为下游地球观测任务的最先进方法。然而,现有的预训练方法通过单一概念视角处理图像,无法捕捉复杂卫星场景的高度组成性。我们提出一种感知组成的预训练框架,该框架显式编码 fractional 土地覆盖混合。每个卫星图像单元被映射到表示其 fractional 土地覆盖分布的直方图,我们将其称为“组成目标”。这些目标作为主要预测目标,并使用 Earth Mover's Distance 提炼到骨干网络中。实验评估表明,感知组成的预训练在需要语义相似性判断的区域级理解任务(包括零样本图像检索和场景分类)上产生显著增益,同时在需要细粒度空间精度的任务(如分割和目标检测)上保持竞争力。使用参数规模为 3680 万的骨干网络,我们的框架在大多数检索和场景分类设置中优于参数分别为 3.03 亿和 6 亿的 SatMAE 和 Prithvi-EO-2.0。在用于组成判别严格测试平台的细粒度 ForestNet-12 数据集上,我们的方法将基线 mAP@10 从 0.279 提升至 0.434,相对提升 55.6%,为显式组成建模的有效性提供直接证据。代码实现可在该 https URL 获取。
英文摘要
Geospatial foundation models have emerged as state-of-the-art methods for downstream Earth observation tasks. However, existing pretraining methodologies process imagery through a single-concept lens, failing to capture the highly compositional nature of complex satellite scenes. We propose a composition-aware pretraining framework that explicitly encodes fractional land-cover mixtures. Each satellite image cell is mapped to a histogram representing its fractional land-cover distribution, which we term the "composition target". These targets serve as the primary prediction objective and are distilled into the backbone using Earth Mover's Distance. Experimental evaluation shows that composition-aware pretraining yields substantial gains on region-level understanding tasks requiring semantic similarity judgment, including zero-shot image retrieval and scene classification, while remaining competitive on tasks requiring fine-grained spatial precision, such as segmentation and object detection. With a 36.8M-parameter backbone, our framework outperforms SatMAE and Prithvi-EO-2.0, which contain 303M and 600M parameters, respectively, in most retrieval and scene classification settings. On the fine-grained ForestNet-12 dataset, a rigorous testbed for compositional discrimination, our method boosts baseline mAP@10 from 0.279 to 0.434, a 55.6% relative improvement, providing direct evidence for the effectiveness of explicit composition modeling. The code implementation can be found at https://github.com/05kashyap/GFM_Composition_Pretraining