发表机构
Space Applications Centre, ISRO; Indian Institute of Science Education and Research Bhopal; Indian Institute of Technology Bombay(印度空间研究组织空间应用中心; 印度科学教育与研究学院博帕尔分校; 印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FairRSFM提出生物群系感知基准与去偏框架,揭示遥感基础模型聚合指标掩盖生态区域性能差异,并提供缓解策略以提升生态鲁棒性。
AI 中文摘要
遥感基础模型(RSFMs)通常使用聚合指标进行评估,这可能会掩盖不同生态区域之间的系统性性能差异。我们引入了FairRSFM,一个用于评估RSFMs中生态群体鲁棒性的生物群系感知基准。FairRSFM将来自14个陆地生物群系类别的地理参考样本映射为六个具有生态意义的宏观组,并在统一的冻结骨干评估协议下评估模型。该基准涵盖四个下游数据集:m-EuroSAT、m-BigEarthNet、m-SA-Crop-Type和带有Dynamic World标签图的MMEarth20K。使用Prithvi-EO-2.0、SatMAE和DOFA在三个随机种子下,我们表明聚合性能始终掩盖了跨架构和任务中依赖生物群系的差异。例如,Prithvi-EO-2.0在m-EuroSAT上达到90.98%的总体宏F1,但平均最差组得分仅为83.72%,而m-SA-Crop-Type的总体mIoU从27.30%下降到Xeric和Mineralogical组中的18.47%。我们进一步评估了生物群系正交线性探测(BOLP)、动态生物群系重加权(DBR)和GroupDRO作为互补的缓解基线。它们的效果依赖于模型和任务;例如,BOLP在不更新RSFM骨干的情况下,将Prithvi-EO-2.0在m-BigEarthNet上的最差组F1@opt从46.12%提高到50.27%。FairRSFM为诊断和缓解遥感基础模型中的生态鲁棒性差距提供了一个可复用的协议。代码和数据集可在以下网址获取:此https URL。
英文摘要
Remote sensing foundation models (RSFMs) are commonly evaluated using aggregate metrics, which can hide systematic performance disparities across ecological regions. We introduce FairRSFM, a biome-aware benchmark for evaluating ecological group robustness in RSFMs. FairRSFM maps georeferenced samples from 14 terrestrial biome classes into six ecologically meaningful macro-groups and evaluates models under a unified frozen-backbone evaluation protocol. The benchmark covers four downstream datasets: m-EuroSAT, m-BigEarthNet, m-SA-Crop-Type, and MMEarth20K with Dynamic World label maps. Using Prithvi-EO-2.0, SatMAE, and DOFA across three random seeds, we show that aggregate performance consistently masks biome-dependent disparities across architectures and tasks. For example, Prithvi-EO-2.0 reaches 90.98% overall macro-F1 on m-EuroSAT but a mean worst-group score of only 83.72%, while m-SA-Crop-Type drops from 27.30% overall mIoU to 18.47% in the Xeric and Mineralogical group. We further evaluate Biome-Orthogonal Linear Probing (BOLP), Dynamic Biome Reweighting (DBR), and GroupDRO as complementary mitigation baselines. Their effectiveness is model- and task-dependent; for example, BOLP improves Prithvi-EO-2.0 worst-group F1@opt on m-BigEarthNet from 46.12% to 50.27% without updating the RSFM backbone. FairRSFM provides a reusable protocol for diagnosing and mitigating ecological robustness gaps in remote sensing foundation models. Code and datasets are available at: https://github.com/aminurhossain/FairRSFM.
Comments13