AI 中文总结
针对水稻叶病分类跨数据集性能下降问题,通过多数据集基准测试,发现强增强有效提升泛化,而自适应批归一化有害,并排除了架构偏置为主要原因。
AI 中文摘要
水稻叶病分类中的跨数据集迁移仍是一个重大挑战,在一个图像集上训练的模型在另一个数据集上部署时性能会显著下降。我们在三个孟加拉水稻叶病数据集(5,419张图像、6个迁移对、3个CNN骨干网络、3个随机种子)上进行了系统性基准测试,以表征和诊断这一失败。强增强在跨数据集宏F1上平均提升了+0.070(Wilcoxon p < 0.001,18个迁移对中有15个为正)。通过分割去除非叶片图像内容显示出方向性益处(平均+0.066,p = 0.062,n = 36个配对观测),该益处在两种独立的分割方法中一致,但未达到常规显著性。自监督ViT对照(DINOv2线性探针)表现出与CNN相当的跨数据集崩溃,排除了架构归纳偏置作为主要驱动因素,并指向采集条件偏移。自适应批归一化(AdaBN)一致地损害迁移性能,且损害程度与源-目标标签先验差异和模型深度相关(Spearman rho = 0.621,p = 0.009)。对12个采样预测的Grad-CAM归因分析无法区分正确与错误的跨域预测(p = 0.462),表明常见的归因代理在实用样本量下不足以诊断偏移。我们在一个具有SHA-256完整性验证的公共仓库中记录了所有冻结结果、预分析标准和可复现性工件。这项工作为理解农业计算机视觉中的跨数据集泛化建立了严格的实证基线,并识别了有效(增强)和无效(AdaBN)的适应策略。
英文摘要
Cross-dataset transfer in rice leaf disease classification remains a significant challenge, with models trained on one image collection performing substantially worse when deployed on another. We conduct a systematic benchmark across three Bangladeshi rice leaf disease datasets (5,419 images, 6 transfer pairs, 3 CNN backbones, 3 random seeds) to characterize and diagnose this failure. Strong augmentation recovers a mean cross-dataset macro-F1 improvement of +0.070 (Wilcoxon p < 0.001, 15 of 18 transfer pairs positive). Removing non-leaf image content via segmentation shows directional benefit (mean +0.066, p = 0.062, n = 36 paired observations) that is consistent across two independent segmentation methods but does not reach conventional significance. A self-supervised ViT control (DINOv2 linear probe) exhibits equivalent cross-dataset collapse to CNNs, ruling out architecture inductive bias as the primary driver and pointing to acquisition-condition shift. Adaptive batch normalization uniformly harms transfer performance, with harm magnitude correlating with source-target label-prior divergence and model depth (Spearman rho = 0.621, p = 0.009). Grad-CAM attribution analysis on 12 sampled predictions does not distinguish correct from incorrect cross-domain predictions (p = 0.462), indicating that common attribution proxies are insufficient for diagnosing shift at practical sample sizes. We document all frozen results, prespecified analysis criteria, and reproducibility artifacts in a public repository with SHA-256 integrity verification. This work establishes a rigorous empirical baseline for understanding cross-dataset generalization in agricultural computer vision and identifies both effective (augmentation) and ineffective (AdaBN) adaptation strategies.
Comments16 pages, 6 figures, 7 tables