发表机构
Turba Open Lab(Turba开放实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对肥料推荐系统难以复用的问题,提出三层开源 Turba 栈,通过版本化快照和离线替代模型,在摩洛哥实现可复现的作物特定推荐,并完成大规模基准测试。
AI 中文摘要
基于位置的肥料推荐系统根据地点、土壤性质、作物类型和生产目标调整养分建议,但当推荐功能主要通过交互界面访问、输出结果未版本化、且训练好的近似模型无法独立加载或进行基准测试时,科学复用受到限制。本技术报告介绍了 Turba 肥料机器学习栈,这是一个用于在摩洛哥实现可复现的基于位置肥料推荐的三层开源实现。turba-client 提供对公开可访问的地点档案、作物特定目标产量空间以及 N、P$_2$O$_5$ 和 K$_2$O 推荐工作流的程序化访问;turba-data 分发分析就绪的快照;turba-models 打包作物特定的推荐输出机器学习替代模型。该架构连接了上游检索、版本化的分析快照、可复现的跨模型基准测试以及可加载的离线替代模型,同时保留了推荐系统输出、观测农业数据和模型生成预测之间的区别。第一个数据集由 44,096 个独特的 ESA WorldCereal 地点构建而成。在中等目标产量设置下,对支持的谷物工作流进行场景扩展,生成了 132,017 个作物-地点推荐请求。由此产生的包含 22 个变量的数据集覆盖 10 个地区、66 个省和 1,149 个市镇。在固定的确定性 80/20 协议下评估了九个回归家族,当前版本打包了五个性能最佳的作物特定模型。机器学习任务是推荐函数仿真,而非预测观测到的作物响应。该栈为空间和时间验证、不确定性估计、田间试验比较以及未来与更多数据的集成提供了可复现的基础。
英文摘要
Site-specific fertilizer recommendation systems adapt nutrient advice to location, soil properties, crop type, and production targets, but scientific reuse is constrained when recommendation functions remain accessible mainly through interactive interfaces, outputs are not versioned, and trained approximations cannot be independently loaded or benchmarked. This technical report presents the Turba fertilizer machine learning stack, a three-layer open-source implementation for reproducible site-specific fertilizer recommendation in Morocco. \texttt{turba-client} provides programmatic access to publicly accessible site profiles, crop-specific target-yield spaces, and N, P$_2$O$_5$, and K$_2$O recommendation workflows; \texttt{turba-data} distributes analysis-ready snapshots; and \texttt{turba-models} packages crop-specific machine learning surrogates of recommendation outputs. The architecture links upstream retrieval, versioned analytical snapshots, reproducible cross-model benchmarking, and loadable offline surrogates while preserving the distinction between recommendation-system outputs, observed agricultural data, and model-generated predictions. The first dataset was constructed from 44,096 unique ESA WorldCereal locations. Scenario expansion across supported cereal workflows generated 132,017 crop-location recommendation requests under a medium target-yield setting. The resulting 22-variable dataset spans 10 regions, 66 provinces, and 1,149 communes. Nine regression families were evaluated under a fixed deterministic 80/20 protocol, and the current release packages five best-performing crop-specific models. The machine learning task is recommendation-function emulation rather than prediction of observed crop response. The stack provides a reproducible basis for spatial and temporal validation, uncertainty estimation, field-trial comparison, and future integration with additional data.