发表机构
Bangladesh Army University of Science & Technology; Rajshahi University of Engineering & Technology(孟加拉国陆军科技大学; 拉杰沙希工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出轻量级卷积网络AgroVisNet和专家验证基准BD-PlantDX,用于孟加拉国萝卜、马铃薯和蛇瓜病害分类,以29万参数实现99.52%准确率,显著优于预训练轻量级模型。
AI 中文摘要
在农艺专业知识稀缺且网络连接不可靠的地区,自动化植物病害诊断正越来越多地部署在农民持有的设备上。三个障碍限制了其实际价值:公开基准数据集主要由少数非本地作物主导,特定区域的数据集很少经过领域专家验证,以及达到有竞争力精度的架构所携带的参数预算不适合低成本硬件。我们提出了AgroVisNet,一个从头训练的紧凑卷积网络,以及BD-PlantDX,一个专家验证的基准数据集,包含12,432张田间图像,涵盖孟加拉国Bogura和Nilphamari地区的萝卜、马铃薯和蛇瓜的12类健康和患病状态。AgroVisNet将带有顺序通道和空间注意力的分组瓶颈残差块与多尺度深度可分离块以及双池化分类头相结合,达到290,572个可训练参数。在BD-PlantDX上,该模型达到99.52%的测试准确率和99.52%的加权F1分数,在相同协议下评估时,超过了所有六个ImageNet预训练的轻量级骨干网络,同时使用的参数少8.7到16.8倍,乘加运算少1.3到8.5倍。导出用于部署时,模型量化为0.46 MB的全整数网络,准确率损失0.22个百分点,并在单个CPU上以8.40毫秒对图像进行分类。在五个随机种子下,准确率保持在99.57 ± 0.10%,十种变体的消融实验隔离了每个组件的贡献,并且相同的架构无需重新设计即可迁移到两个独立收集的数据集,准确率分别为98.71%和99.05%。Grad-CAM证据表明,预测基于带有病斑的叶片区域而非背景线索。
英文摘要
Automated plant disease diagnosis is increasingly deployed on farmer-held devices in regions where agronomic expertise is scarce and network connectivity is unreliable. Three obstacles limit its practical value: public benchmarks are dominated by a small set of non-native crops, region-specific datasets are rarely validated by domain experts, and the architectures that reach competitive accuracy carry parameter budgets that are unsuited to low-cost hardware. We propose AgroVisNet, a compact convolutional network trained from scratch, together with BD-PlantDX, an expert-validated benchmark of 12,432 field images spanning 12 classes of radish, potato and pointed gourd in healthy and diseased states, collected across the Bogura and Nilphamari districts of Bangladesh. AgroVisNet couples grouped bottleneck residual blocks carrying sequential channel and spatial attention with multi-scale depthwise blocks and a dual-pooling classification head, reaching 290,572 trainable parameters. On BD-PlantDX the model attains 99.52% test accuracy and 99.52% weighted F1, exceeding all six ImageNet-pretrained lightweight backbones evaluated under an identical protocol while using 8.7 to 16.8 times fewer parameters and 1.3 to 8.5 times fewer multiply-accumulate operations. Exported for deployment, the model quantises to a 0.46 MB full-integer network at a 0.22 percentage-point accuracy cost and classifies an image in 8.40 ms on a single CPU. Across five random seeds accuracy remains at 99.57 +- 0.10%, a ten-variant ablation isolates the contribution of each component, and the same architecture transfers without redesign to two independently collected datasets at 98.71% and 99.05% accuracy. Grad-CAM evidence indicates that predictions rest on lesion-bearing leaf regions rather than on background cues.