arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Infra-Bench CLS:面向关键基础设施分类的地球观测基础模型的全球开源基准

Infra-Bench CLS: A Global, Open-Source Benchmark for Critical Infrastructure Classification with Earth Observation Foundation Models

Justin Guthrie, Edward Oughton, Konrad Wessels, Matthew Rice, Isaac Corley

arXiv 2609.09482首次发表:更新:

发表机构

George Mason University; Taylor Geospatial Institute(乔治梅森大学; 泰勒地理空间研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出 Infra-Bench CLS 基准,评估七个地球观测基础模型在关键基础设施分类上的表现,最佳模型宏 F1 达 57.9%,较基线提升 48%,但电力类表现欠佳。

AI 中文摘要

关键基础设施的位置数据在全球范围内往往不完整且分布不均,尤其是在发展中地区。地球观测基础模型被视为使我们能够更有效地理解自然和建筑环境的新步骤,这引发了关于它们在执行具有挑战性的下游任务时的有效性的问题。然而,基础模型在检测和分类设施规模的关键基础设施方面仍未得到充分测试,而这些基础设施支撑着一系列重要的社会和经济功能。因此,我们引入了 Infra-Bench CLS 作为基准,用于在覆盖七大洲和 13 个基础设施类别的 18,756 张 Sentinel-1 SAR 和 Sentinel-2 多光谱设施规模关键基础设施资产图像上测试基础模型,并报告了保留的 10 个类别的结果。通过使用线性探测和微调,针对两个训练数据集级别(1.0 倍和 0.3 倍),评估了七个基础模型(SatlasPretrain S2、SatlasPretrain S1、CROMA、Prithvi-EO-2.0、AlphaEarth Foundations、OlmoEarth v1.1-Base 和 DINOv3 ViT-L/16)。当将宏 F1 分数与 ResNet-18 监督基线(39.2%)进行比较时,最佳基础模型达到了 57.9%,提升了 48%。表现最好的类别是机场(F1 85.3%)、火车站(F1 82.1%)和数据中心(F1 77.6%)。相比之下,许多电力行业类别表现不佳(F1 27.5-46.2%)。这些发现表明,基础模型可以实现更优的关键基础设施分类,但未来的工作应评估更高分辨率图像上的性能,特别是对于表现不佳的行业,如电力。

英文摘要

Critical infrastructure location data is often incomplete and unevenly distributed globally, especially in developing regions. Earth observation foundation models are proposed as a new step in enabling us to more efficiently understand the natural and built environment, raising questions as to their effectiveness in performing challenging downstream tasks. Yet, foundation models remain largely untested for detecting and classifying the facility-scale critical infrastructure that underpins a range of important societal and economic functions. Subsequently, Infra-Bench CLS is introduced as a benchmark to test foundation models on 18,756 Sentinel-1 SAR and Sentinel-2 multispectral facility-scale critical infrastructure asset images covering seven continents and 13 infrastructure classes, with results reported for the 10 retained classes. Using linear probing and fine-tuning for two training dataset levels (1.0x and 0.3x), seven foundation models are evaluated (SatlasPretrain S2, SatlasPretrain S1, CROMA, Prithvi-EO-2.0, AlphaEarth Foundations, OlmoEarth v1.1-Base, and DINOv3 ViT-L/16). When comparing macro F1 scores to a ResNet-18 supervised baseline of 39.2 percent, the best foundation model achieved 57.9 percent, a 48 percent improvement. Top performing classes were airports (F1 85.3 percent), train stations (F1 82.1 percent), and data centers (F1 77.6 percent). By contrast, many of the power sector classes perform poorly (F1 27.5-46.2 percent). These findings suggest foundation models can enable superior critical infrastructure classification, but future work should evaluate performance on higher-resolution imagery, particularly for poorly performing sectors, such as power.

Comments9 figures. Supporting information with 10 figures and 13 tables. Submitted to Big Earth Data

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑