基于openEO的地球观测数据立方体机器学习API
A Machine Learning API for Earth Observation Data Cubes Based on openEO
- Institute for Geoinformatics, University of Münster(明斯特大学地理信息学研究所)
- FGV Agro - Center for Agribusiness Studies, Fundação Getulio Vargas(热图利奥·瓦加斯基金会农业研究中心)
- Bochum University of Applied Sciences(波鸿应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对EO数据立方体与ML方法不匹配问题,提出openEO流程级ML规范,分三阶段支持经典与深度学习,原型验证跨后端互操作性,提升可复现性与可移植性。
中文摘要 AI 辅助
地球观测(EO)数据日益组织为时空数据立方体,而机器学习(ML)方法则基于表格特征矩阵或结构化张量输入进行操作。这种不匹配迫使进行特定于平台的转换,这些转换难以在云基础设施之间复现或迁移。openEO规范为跨异构后端的地球观测数据访问和处理提供了统一接口,但缺乏标准化的机器学习集成方法。我们提出了一种面向openEO的流程级机器学习规范,分为三个阶段:模型初始化、模型操作(训练、调优、推理、验证)和模型管理。该规范支持随机森林和支持向量机等经典算法,以及用于时间序列和空间斑块建模的深度学习架构,包括TempCNN、时间注意力编码器和基础模型推理。R和Python中的三个原型实现展示了跨多种技术栈的可行性。一个作物类型制图用例通过向独立的R和Python后端提交相同的流程图并比较预测和评估指标,展示了跨后端互操作性。另外两个用例展示了时间序列上的深度学习和基础模型推理,每个用例在专用后端上执行。然而,原型揭示,完全的跨后端可移植性需要对序列化格式和执行语义进行比流程级更深入的协调;规范边界之外的后端库版本和预处理约定也会影响可复现性。通过显式的后端一致性配置文件解决这两个问题是近期最重要的方向。该规范促进了云平台上地球观测数据立方体上机器学习工作流的可复现性、可移植性和可访问性。
英文摘要
Earth Observation (EO) data are increasingly organized as spatio-temporal data cubes, while machine learning (ML) methods operate on tabular feature matrices or structured tensor inputs. This mismatch forces platform-specific transformations that are difficult to reproduce or transfer across cloud infrastructures. The openEO specification provides a unified interface for EO data access and processing across heterogeneous backends, but lacks a standardized approach for ML integration. We propose a process-level ML specification for openEO structured into three stages: model initialization, model actions (training, tuning, inference, validation), and model management. It supports classical algorithms such as Random Forest and SVM, as well as deep learning architectures for time-series and spatial patch-based modeling, including TempCNN, Temporal Attention Encoders, and foundation model inference. Three prototype implementations in R and Python demonstrate feasibility across diverse technology stacks. A crop type mapping use case demonstrates cross-backend interoperability by submitting an identical process graph to independent R and Python backends and comparing predictions and evaluation metrics. Two further use cases demonstrate deep learning on time series and foundation model inference, each executed on a dedicated backend. The prototypes reveal, however, that full cross-backend portability requires deeper harmonization of serialization formats and execution semantics than the process level alone can enforce; backend library versions and preprocessing conventions outside the specification's boundary also affect reproducibility. Addressing both through explicit backend conformance profiles represents the most important near-term direction. The specification advances the reproducibility, portability, and accessibility of ML workflows on EO data cubes across cloud platforms.