发表机构
University of Alabama in Huntsville; NASA Marshall Space Flight Center(阿拉巴马大学亨茨维尔分校; NASA马歇尔太空飞行中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究评估Prithvi作物分类基础模型在三大洲的分布外性能,发现物候错位导致精度下降,通过类别合并和窗口压缩可恢复精度,为模型可靠部署提供指导。
AI 中文摘要
基于大型卫星档案预训练并经过微调的地理空间基础模型(GeoFMs)已被证明能提高作物分类精度和地理可迁移性。然而,它们在训练分布之外的运行性能仍缺乏充分表征。我们评估了广泛采用的GeoFM [Prithvi-EO-2.0]在三大洲12个国家37个事件中的分布外性能,并与区域参考产品进行了验证。结果表明,平均总体精度(OA)从美国的0.65下降到欧洲的0.40。除精度指标外,我们评估了模型性能的五个关键方面:模型置信度是否指示信号失败、对观测窗口的敏感性、粗化类别方案的影响,以及对波段丢失和云影污染的鲁棒性。当观测窗口与当地作物物候错位时,精度崩溃,而确定性置信度仍然很高。八对事件中有七对的期望校准误差增加,每个受影响场景的12%-51%被以接近零的精度高置信度错误标记。蒙特卡洛dropout熵在所有八个事件中均记录了偏移,表明跨大陆表观下降的大部分反映了物候错位而非空间迁移。两项调整在不重新训练的情况下恢复了精度。基于模型的主要混淆将13个类别合并为10个,使平均OA提高了8.4个百分点。将窗口压缩至近实时使用,在45至90天的平台期内保持了精度,峰值出现在约75天,而严于30天的窗口比该平台期低约0.11。因此,微调的作物GeoFM仅在观测窗口与当地生长季节匹配时才有效迁移。我们将这些发现转化为运行指南,以支持所发布模型的可靠部署。
英文摘要
Fine-tuned geospatial foundation models (GeoFMs) pretrained on large satellite archives have been shown to improve crop classification accuracy and geographic transferability. However, their operational performance beyond the training distribution remains poorly characterized. We evaluated the out-of-distribution performance of a widely adopted GeoFM [Prithvi-EO-2.0] across 37 events in 12 countries on three continents and validated against regional reference products. Results indicated that the mean overall accuracy (OA) declined from 0.65 in the United States to 0.40 in Europe. Beyond accuracy metrics, we assessed five key aspects of model performance: whether model confidence indicates signal failure, sensitivity to observation windows, the effect of coarsening class schemes, and robustness to both band loss and cloud- and shadow-contamination. Accuracy collapsed when the observation window misaligned with local crop phenology, while deterministic confidence remained high. Expected calibration error increased for seven of eight paired events, and 12-51% of each affected scene was confidently mislabeled at near-zero precision. Monte Carlo dropout entropy registered the shift in all eight, indicating that much of the apparent cross-continent decline reflected phenological misalignment rather than spatial transfer. Two adjustments recovered accuracy without retraining. Consolidating 13 classes into 10, based on the model's dominant confusions, raised the mean OA by 8.4 percentage points. Compressing the window toward near-real-time use preserved accuracy across a 45- to 90-day plateau, peaking near 75 days, though arms tighter than 30 days fell about 0.11 below that plateau. Fine-tuned crop GeoFMs therefore transfer usefully only where observation windows match local growing seasons. We translate these findings into operational guidance for the reliable deployment of the released model.