arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgroBench:一个可复现的多模态基准,用于从县级统计数据和像素观测中进行弱监督作物产量学习

AgroBench: A Reproducible Multimodal Benchmark for Weakly Supervised Crop Yield Learning from County Statistics and Pixel Observations

Udaiveer Singh, Rajiv Ranjan, Shashank Tamaskar, Dharmendra Saraswat

arXiv 2609.26809首次发表:更新:

发表机构

Plaksha University; Purdue University(普拉克沙大学; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对农业产量统计尺度粗与像素级监督需求不匹配的问题,提出可复现基准AgroBench,将县级产量转化为弱监督像素级时间序列,并建立评估协议,为多模态遥感与弱监督学习提供基础。

AI 中文摘要

可靠的农业产量统计数据通常以较粗的行政尺度报告,而现代地理空间机器学习方法需要空间显式的、像素级别的监督。这种不匹配限制了利用多模态地球观测数据进行作物产量学习的大规模基准的发展。本文提出了一个可复现的基准AgroBench,用于将公开可用的美国县级作物产量统计数据转化为弱监督的像素级作物时间序列。每个作物像素时间序列与一个县级产量值配对,作为弱监督信号,而非直接测量的像素级产量标签。我们的地理空间数据生成流程整合了美国农业部作物产量统计数据、作物特定土地覆盖掩膜、Sentinel-2多光谱图像、Sentinel-1合成孔径雷达观测、气候变量和地形信息,以生成描述整个生长季节中单个作物像素的时间对齐的多模态序列。由此产生的基准包含超过1300万个观测数据,来自788,654个独特作物像素,涵盖5,107个县年组合,跨越八个生长季节(2017年至2024年),涉及五种美国主要作物。为了促进标准化评估,我们使用留一年法评估协议建立了作物产量预测基准,并提供了使用代表性机器学习模型的基线结果。通过发布完整的数据生成流程、基准数据集和评估协议,AgroBench为弱监督学习、多模态遥感、时空建模和农业地理空间基础模型的未来研究提供了可复现的基础。

英文摘要

Reliable agricultural yield statistics are typically reported at coarse administrative scales, whereas modern geospatial machine learning methods require spatially explicit, pixel level supervision. This mismatch has limited the development of large-scale benchmarks for crop yield learning using multimodal Earth observation data. A reproducible benchmark, AgroBench, is presented for transforming publicly available U.S. county level crop yield statistics into weakly supervised pixel-level crop time series. Each crop pixel time series is paired with a county-level yield value as a weak supervisory signal rather than a directly measured pixel-level yield label. Our geospatial data generation pipeline integrates USDA crop yield statistics with crop-specific land cover masks, Sentinel 2 multispectral imagery, Sentinel-1 synthetic aperture radar observations, climatic variables, and terrain information to produce temporally aligned multimodal sequences describing individual crop pixels throughout the growing season. The resulting benchmark contains over 13 million observations from 788,654 unique crop pixels spanning 5,107 county year combinations across eight growing seasons (2017 to 2024) for five major U.S. crops. To facilitate standardized evaluation, we establish a crop yield prediction benchmark using a Leave-One-Year-Out evaluation protocol and provide baseline results using representative machine learning models. By releasing the complete data generation pipeline, benchmark dataset, and evaluation protocol, AgroBench provides a reproducible foundation for future research in weakly supervised learning, multimodal remote sensing, spatiotemporal modeling, and geospatial foundation models for agriculture.

Comments13 Pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑