发表机构
Korea University; FPT University(高丽大学; FPT大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SOLAR多模态生成模型,通过融合专家模块对齐视觉与语言特征,将番茄病害分析转化为生成式视觉问答任务,在41,677张图像上优于现有模型,实现多任务诊断理解。
AI 中文摘要
植物病害分析的人工智能已经从特定任务分类器发展到能够联合解释视觉和文本信息的多模态模型。然而,在实际精准农业部署中仍存在局限性,因为大多数现有方法将病害理解视为孤立的预测任务,未能捕捉症状识别、严重程度评估和基于问题的诊断推理之间的互补关系。在番茄病理学中,对病害叶片的准确解释不仅需要标签预测,还需要将视觉症状与语义上下文整合,以支持全面且可解释的理解。在此,我们提出SOLAR,一种多模态生成模型,可理解跨越六个问答任务的番茄病害。SOLAR通过基于专家混合的融合专家模块,学习将视觉特征与任务感知的语言表示对齐,使其能够在各种诊断任务中生成上下文相关的答案。通过将番茄病害分析表述为生成式视觉问答(VQA)任务,SOLAR提供了一个灵活的框架,支持在单一模型内进行多任务推理,同时提高性能和跨任务知识共享。我们在41,677张图像上评估SOLAR,包括216,209个问答(QA)对,以在封闭式和开放式QA设置下理解番茄叶部病害。实验结果表明,SOLAR在所有任务中始终优于最先进的纯视觉、视觉-语言和特定任务模型,展现出卓越的准确性、鲁棒性和多模态推理能力。这些发现凸显了生成式多模态建模作为植物病害理解有效方向的潜力。本研究的代码可在以下网址获取:https URL。
英文摘要
Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical deployment in precision agriculture remains limited because most existing approaches treat disease understanding as isolated prediction tasks, failing to capture the complementary relationships among symptom recognition, severity assessment, and question-driven diagnostic reasoning. In tomato pathology, accurate interpretation of diseased leaves requires more than label prediction; it demands integrating visual symptoms with semantic context to support a comprehensive and explainable understanding. Here, we present SOLAR, a multimodal generative model that understands tomato disease spanning six question-answering tasks. SOLAR learns to align visual features with task-aware language representations by Fusion Expert module based on mixture-of-expert, enabling it to generate contextually relevant answers across diverse diagnostic tasks. By formulating tomato disease analysis as a generative Visual Question Answering (VQA) task, SOLAR provides a flexible framework that supports multi-task inference within a single model while improving performance and cross-task knowledge sharing. We evaluate SOLAR on $41,677$ images, including $216,209$ Question-Answering (QA) pairs to understand tomato leaf disease under both closed and open-ended QA settings. Experimental results show that SOLAR consistently outperforms state-of-the-art vision-only, vision-language, and task-specific models across all tasks, demonstrating superior accuracy, robustness, and multimodal reasoning. These findings highlight the potential of generative multimodal modeling as an effective direction for understanding of plant disease. The code for this study is available at https://github.com/EnalisUs/SOLAR.
CommentsIn submission to Computers and Electronics in Agriculture Journal