CoAtNet-DeepMoE:一种结合DeepSeek混合专家模型的卷积-注意力混合架构,用于参数高效的番茄病害分类
CoAtNet-DeepMoE: A Convolution-Attention Hybrid with DeepSeek Mixture-of-Experts for Parameter-Efficient Tomato Disease Classification
- Begum Rokeya University(贝古姆罗基亚大学)
- Texas State University(德克萨斯州立大学)
- Asian Institute of Technology(亚洲理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对番茄病害分类中模型参数多与性能下降的矛盾,提出CoAtNet-DeepMoE混合架构,结合DeepSeek混合专家减少参数,在Kaggle和PlantVillage上以2.47M参数达到约99.8%的SOTA性能。
AI中文摘要:
世界人口正在快速增长,技术也在同步进步。满足这70亿人对粮食的巨大需求,不仅依赖于增加粮食产量,还依赖于减少粮食损失。由病害造成的作物损失既影响粮食供应,也影响一个国家的财政和经济稳定。番茄是全球最主要的粮食作物之一,而其中相当大一部分产量因病害而损失。人们曾使用机器学习技术进行番茄病害的特征提取和早期诊断,如今,基于深度学习模型的方法被广泛用于病害识别。然而,现有的大多数模型参数密集,这增加了训练和推理所需的时间。因此,虽然轻量级模型更适合用户友好的应用,但它们往往表现出性能下降。为了平衡性能与模型大小,我们提出了CoAtNet-DeepMoE,一种用于丰富特征提取的卷积-注意力混合架构,并进一步通过DeepSeek混合专家模型增强,在不牺牲准确率的情况下大幅减少参数数量。我们在来自Kaggle和PlantVillage的平衡和不平衡数据集上评估了我们的模型,证明了其鲁棒性,并在Kaggle上实现了99.80%的准确率、99.80%的精确率、99.80%的召回率和99.80%的F1分数,在PlantVillage上实现了99.83%的准确率、99.85%的精确率、99.76%的召回率和99.80%的F1分数,代表了仅用2.47M参数的最先进性能。源代码将在该https URL上提供。
英文摘要:
The world population is growing rapidly, and technology is improving in parallel. Meeting the huge demand for food for these 7 billion people not only depends on increasing food production but also on reducing food loss. Crop losses due to disease affect both the food supply and the financial and economic stability of a country. Tomatoes are among the top food-producing crops globally, and a significant portion of this production is lost due to disease. People have used Machine Learning techniques for feature extraction and early diagnosis of tomato diseases, and nowadays, Deep Learning-based models are widely used for disease recognition. However, most existing models are highly parameter-intensive, which increases the time required for training and inference. As a result, while lightweight models are more suitable for user-friendly applications, they often show a reduction in performance. To balance performance and model size, we propose CoAtNet-DeepMoE, a Convolution-Attention hybrid architecture for rich feature extraction, further enhanced with a DeepSeek Mixture of Experts to substantially reduce the number of parameters without sacrificing accuracy. We evaluate our model on both balanced and imbalanced datasets from Kaggle and PlantVillage, demonstrating robustness and achieving 99.80% accuracy, 99.80% precision, 99.80% recall, and 99.80% F1-score on Kaggle, and 99.83% accuracy, 99.85% precision, 99.76% recall, and 99.80% F1-score on PlantVillage, representing state-of-the-art performance with only 2.47M parameters. The source code will be available at https://github.com/nadimbrur/CoAt-MoE.