发表机构
Technical University of Munich (TUM); German Aerospace Center (DLR); Massachusetts Institute of Technology (MIT); Taylor Geospatial(慕尼黑工业大学; 德国航空航天中心; 麻省理工学院; 泰勒地理空间公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出GeoFMs仅用准确率排序的不足,分析其校准度与分布偏移敏感性,发现EO预训练模型更易过度自信,提出需多条件多指标评估以缩小实际部署差距。
AI 中文摘要
地理空间基础模型(GeoFMs)通常通过标准基准条件下的平均排名进行排序和选择。我们表明,该协议过于狭窄:其在关键地球观测(EO)任务中的部署需要进一步的分析角度,主要是校准,即模型的置信度与其正确性之间的一致性。在16个冻结编码器、4个分类数据集和5个分割数据集,以及两个正交压力轴下,随着损坏程度加剧,每个编码器的性能都会下降,且排名也会发生变化。在四个分类基准中,经EO预训练和ImageNet预训练的编码器在干净数据的准确率和校准度上无法区分,且EO预训练在偏移下的稳定性并不比ImageNet预训练更好。在偏移下,GeoFMs比ImageNet预训练编码器更加过度自信,在每个损坏等级和每个损坏类别中均如此。中心核对齐(CKA)分析将此与表示刚性联系起来:经EO预训练的嵌入在损坏下移动较少,同时损失的任务信息与ImageNet预训练的嵌入相当,且仍保持过度自信。我们应用三种常用的不确定性量化方法,发现温度缩放和深度集成无法抵消这种性能下降,而高斯过程探针仅在严重云损坏下将预期校准误差(ECE)减半,却在干净数据上将其加倍。在选择性预测实验中,我们发现基于置信度的弃权(不执行)无法规避自信的错误预测,因此主张基准排名和评估应在多种条件和指标下进行,以更全面地评估模型开发进展,缩小与实际部署场景的差距。
英文摘要
Geospatial Foundation Models (GeoFMs) are most commonly ranked and selected by accuracy on standard benchmark conditions via averaged ranks. We show that this protocol is too narrow: the promised deployment in critical EO tasks requires further angles of analysis, mainly calibration, the agreement between a model's confidence and its correctness. Across 16 frozen encoders, four classification and five segmentation datasets, and two orthogonal stress axes, every encoder degrades as corruption intensifies, and the ranking changes as well. Across the four classification benchmarks, EO-pretrained and ImageNet-pretrained encoders are indistinguishable on clean accuracy and clean calibration, and EO pretraining provides no more stability under shift than ImageNet pretraining. Under shift the GeoFMs drift further into overconfidence than the ImageNet-pretrained encoders, at every grade and in every corruption family. A centered kernel alignment (CKA) analysis ties this to representational rigidity: EO-pretrained embeddings move less under corruption while losing just as much task information and remaining overconfident. We apply three commonly explored uncertainty quantification methods and find that temperature scaling and deep ensembles cannot counteract the degradation, while a Gaussian-process probe roughly halves ECE under severe cloud only by tripling it on clean data. In selective prediction experiments, we find that confidence-based abstention cannot defer around confidently wrong predictions, and advocate that benchmark rankings and evaluations should therefore operate across a multitude of conditions and metrics to more holistically evaluate model development progress and close the gap to real world deployment scenarios.