发表机构
University of Münster; IBM Research; Norwegian Computing Center(明斯特大学; IBM研究院; 挪威计算中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过对TerraMind和THOR两个地理空间基础模型对比,研究其在补丁大小、解码器复杂性等方面的差异轴,发现架构设计选择对性能差异影响更大,体现互补投资策略,还得出假设和诊断消融方法,有望推广到未来模型。
AI 中文摘要
地理空间基础模型(GFM)的基准越来越多地通过综合分数对模型进行排名,但这种排名掩盖了模型差异的原因:差距中有多少是架构造成的,多少是解码器能力造成的,多少是特定用例的产物?本研究通过对欧洲航天局$\Phi$实验室开发的具有不同设计理念的两个GFM进行对比,解决了这一差距:THOR引入了支持可变补丁大小的计算自适应架构,并在其原生分辨率下统一了哨兵1、2和3的数据;TerraMind是一种多模态生成式GFM,通过双尺度令牌/像素目标进行预训练,能够在推理时进行任意到任意的跨模态生成(模态思考)以推断缺失的传感器。我们研究了这两种架构在十个用例中的实际差异轴——补丁大小、解码器复杂性、微调机制、输入模态和模型规模——涵盖不同领域的分割和回归,包括气候灾害响应、甲烷泄漏检测、积雪监测或海冰测绘。我们发现,架构设计选择——特别是补丁大小和解码器类型——比模型本身更能解释性能差异,这两种模型体现了互补的投资策略(TerraMind的预训练时间尺度与THOR的推理时间令牌化),并且正确解释结果需要数据集级别的特征描述。结果不是单一的赢家,而是一组假设和一种诊断消融方法,我们期望将其推广到THOR和TerraMind之外的未来GFM。
英文摘要
Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific artefact? This study addresses that gap through a controlled comparison of two GFMs developed under European Space Agency's $Φ$-lab with contrasting design philosophies: THOR, which introduces a compute-adaptive architecture supporting variable patch sizes and unifies Sentinel-1, -2, and -3 data at their native resolutions; and TerraMind, a multimodal generative GFM pretrained with a dual-scale token/pixel objective that enables any-to-any cross-modal generation (Thinking-in-Modalities) to infer missing sensors at inference time. Rather than reporting a single leaderboard, we investigate the axes along which the two architectures actually differ - patch size, decoder complexity, finetuning regime, input modality, and model scale - across ten use cases spanning segmentation and regression in diverse domains, including climate disaster response, methane leak detection, snow monitoring, or sea ice mapping. We find that architectural design choices - patch size and decoder type in particular - explain more performance variance than model identity itself, that the two models embody complementary investment strategies (pretraining-time scale for TerraMind versus inference-time tokenisation for THOR), and that correctly interpreting results requires dataset-level characterisation. The resulting picture is not a single winner but a set of hypotheses and a diagnostic ablation methodology that we expect to generalise to future GFMs beyond THOR and TerraMind.