发表机构
The Islamia University of Bahawalpur; Shihezi University; Government College of Women University of Faisalabad(巴哈瓦尔布尔伊斯兰大学; 石河子大学; 费萨拉巴德女子政府学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对巴基斯坦旁遮普省小样本产量数据,构建集成学习与叶片健康分类原型,揭示随机验证高R2源于数据重复,按年份留出时模型失效,线性趋势更优,并报告叶片分类与产量无关联的负面结果。
AI 中文摘要
产量预测有助于规划者和农民决定投入、储存和进口,但小规模农业数据表可能使报告的准确性在新季节失效。我们为巴基斯坦旁遮普省构建了一个作物产量原型,该原型结合了随机森林、XGBoost、支持向量回归和Ridge堆叠集成,以及一个MobileNetV2叶片健康分类器,并将其部署为带有SHAP解释的Web应用程序。对合并的Kaggle派生数据表(414行)进行随机80/20划分,集成模型的R2为0.991。审计显示,该表仅包含46个独立观测值:一个连接每年九条温度记录的连接操作将每个作物年份复制了九次。按整年留出验证时,相同行上的XGBoost的R2从0.994降至-0.20。在去重数据以及更长的FAOSTAT数据表(1990-2024年,70个观测值)上,按作物划分的线性趋势(留一年交叉验证的RMSE为0.29吨/公顷)优于所有未提供年份的模型(0.84-1.06)。农药使用和全国温度变化均无法解释趋势残差。Kaggle表中明显的农药收益在FAOSTAT上消失,其中农药序列大部分为插补值,且两个版本数据不一致。一个独立的区级小麦面板(36个区,13个季节)显示,旁遮普省的大部分变异是空间性的,区均值加上共同趋势与学习模型的表现相当。叶片分类器在留出的PlantVillage图像上达到99.87%的准确率,并在去除近重复后对336张未见过的病害叶片召回96.4%。然而,没有叶片图像与产量记录配对,因此产量模型中使用的健康评分不得不构建,且未增加价值。我们连同原型一起报告这些负面发现。
英文摘要
Yield forecasts help planners and farmers decide on inputs, storage and imports, but small agricultural tables can make reported accuracy fail on a new season. We built a crop-yield prototype for Punjab, Pakistan that combines Random Forest, XGBoost, support vector regression and a Ridge-stacked ensemble with a MobileNetV2 leaf-health classifier, and deployed it as a web application with SHAP explanations. A random 80/20 split of a merged Kaggle-derived table (414 rows) gives the ensemble an R2 of 0.991. An audit showed that the table contains only 46 independent observations: a join with nine temperature records per year copied every crop-year nine times. Holding out whole years takes XGBoost on the same rows from R2 = 0.994 to -0.20. On deduplicated data, and on a longer FAOSTAT table (1990-2024, 70 observations), a per-crop linear trend (leave-one-year-out RMSE 0.29 Ton/Ha) beats every model not given the year (0.84-1.06). Neither pesticide use nor national temperature change explains the trend residuals. An apparent pesticide gain in the Kaggle table disappears on FAOSTAT, where the pesticide series is mostly imputed and the two releases disagree. An independent district-level wheat panel (36 districts, 13 seasons) shows that most variation in Punjab is spatial and that a district mean with a common trend matches the learned models. The leaf classifier reaches 99.87% accuracy on held-out PlantVillage images and recalls 96.4% of 336 unseen diseased leaves after near-duplicates were removed. However, no leaf image is paired with a yield record, so the health score used in the yield models had to be constructed and adds nothing. We report these negative findings together with the prototype.