arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08727cs.CVcs.AI

TomaMMU:用于番茄叶片病害的综合多模态理解基准

TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases

  • FPT University(FPT大学)
  • Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

Gia-Han Truong, Khang Nguyen Quoc, Luyl-Da Quach

AI总结:

本研究推出用于番茄叶片病害多模态理解的TomaMMU数据集及TomaBench基准,评估14种VLMs时发现其细粒度识别等存在差距,微调后MCQ准确率达96.09%,为相关研究提供了方向。

AI中文摘要:

为解决这一空白,我们推出了TomaMMU,这是一个大规模的番茄叶片病害多模态理解数据集,同时还推出了TomaBench,这是一个用于评估视觉语言模型(VLMs)在番茄病害理解方面性能的基准。TomaMMU包含28808张高质量图像,涵盖15个类别,以及213119个人工标注的视觉问答对,这些问答对是通过包含数据收集、人工标注和问答生成三个阶段的流程生成的。基于此基础,TomaBench将7项农业任务组织成一个分层三级分类体系,涵盖基础感知、病理学理解和专家诊断,共同实现从低级视觉识别到高级诊断推理的系统评估。这些任务评估视觉症状识别、分类关系和诊断推理,全面展现模型对植物病理学的掌握程度。我们对14种最先进的VLMs的研究结果显示,它们在细粒度识别和基于事实的推理方面存在显著差距,在具有挑战性的多项选择题(MCQs)和开放式问题上的表现始终不佳。这些结果表明,当前的VLMs难以将视觉感知转化为可靠的诊断知识,因此需要针对性的领域适配。在TomaMMU上进行简单的微调就大幅缩小了这一差距,使具有挑战性的MCQs的准确率提升至96.09%,优于近期的VLMs,为未来工作指明了有前景的方向。所有数据和代码可在该https URL获取。

英文摘要:

To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning 15 categories and 213,119 human-annotated visual question-answer pairs, generated through a three-stage pipeline comprising Data Collection, Human Annotation, and Question-Answer Generation. Building on this foundation, TomaBench organizes seven agricultural tasks into a hierarchical three-level taxonomy spanning Basic Perception, Pathology Understanding, and Expert Diagnosis, which together enable systematic evaluation from low-level visual recognition to high-level diagnostic reasoning. The tasks assess visual symptom recognition, taxonomic relationships, and diagnostic reasoning, offering a comprehensive view of how well models grasp plant pathology. Our results pronounced gaps in fine-grained recognition and factually grounded reasoning with 14 state-of-the-art VLMs, consistently underperforming on both challenging MCQs and open-ended questions. These results suggest that current VLMs struggle to translate visual perception into reliable diagnostic knowledge, motivating the need for targeted domain adaptation. Simple fine-tuning on TomaMMU substantially narrows this gap, boosting accuracy on challenging MCQs to 96.09%, outperforming recent VLMs, and pointing toward promising directions for future work. All data and code is available in https://huggingface.co/datasets/enalis/TomaMMU.

补充信息

↑