AgroGround:农业中的多粒度接地识别
AgroGround: Multi-Granularity Grounded Recognition in Agriculture
浏览论文内容
中文总结 AI 辅助
AgroGround提出大规模接地农业识别数据集,通过自动化标注和联合识别定位指令微调,显著提升模型在身份、区域及健康图像弃权上的性能,并建立基准。
中文摘要 AI 辅助
农业视觉模型通常被评估为仅进行识别或定位,但可靠的诊断需要识别存在的内容并定位证据。农业视觉问答(VQA)数据集携带丰富的语义标签,但很少将它们与图像区域关联起来,而手动添加此类注释在大规模上成本高昂。我们引入了AgroGround,一个用于接地农业识别的大规模数据集:识别植物病害和其他农业目标并定位其图像区域。一个自动化流程将八个农业VQA数据集的标签转换为病害病变和整个对象的注释,产生了794,850个指令示例。健康图像为病害查询提供负监督,教导模型返回空预测。我们在已知目标接地指令与需要识别和定位的指令结合上微调了一个共享的视觉语言模型。我们在与所有训练数据不相交的1,480张人工验证图像上评估了预测身份、区域、联合正确性和健康图像弃权(不执行)。仅接地微调将识别准确率从51.8%降至29.1%,而添加识别和定位指令将其提升至72.6%。在图像和注释固定的情况下,结合两种格式将联合准确率从19.2%提高到43.3%,同时接地性能相当。健康负样本将健康图像上的弃权(不执行)率提高到95.0%,强化学习改善了病变级接地。生成的2B模型在接地F1分数上超过了其注释教师,在我们的基准和外部PlantSeg测试集上均如此。AgroGround为接地农业识别建立了一个基准,衡量身份和定位的联合正确性以及健康图像上的弃权(不执行)。代码可在以下https URL获取。
英文摘要
Agricultural visual models are typically evaluated for either recognition or localization, but reliable diagnosis requires identifying what is present and localizing the evidence. Agricultural visual question answering (VQA) datasets carry rich semantic labels but rarely link them to image regions, and adding such annotations by hand is costly at scale. We introduce AgroGround, a large-scale dataset for grounded agricultural recognition: identifying plant diseases and other agricultural targets and localizing their image regions. An automated pipeline converts the labels of eight agricultural VQA datasets into annotations for disease lesions and whole objects, producing 794,850 instruction examples. Healthy images provide negative supervision for disease queries, teaching the model to return empty predictions. We fine-tune a shared vision-language model on known-target grounding instructions combined with instructions requiring both recognition and localization. We evaluate predicted identities, regions, joint correctness, and healthy-image abstention on 1,480 human-verified images disjoint from all training data. Grounding-only fine-tuning reduces recognition accuracy from 51.8\% to 29.1\%, while adding recognition-and-localization instructions raises it to 72.6\%. With images and annotations held fixed, combining the two formats raises joint accuracy from 19.2\% to 43.3\% at comparable grounding. Healthy negatives raise abstention on healthy images to 95.0\%, and reinforcement learning improves lesion-level grounding. The resulting 2B model exceeds its annotation teacher in grounding F1 on our benchmark and on the external PlantSeg test set. AgroGround establishes a benchmark for grounded agricultural recognition, measuring joint correctness of identity and localization along with abstention on healthy images. The code is available at https://github.com/AB-Abdulla/AgroGround.
发表机构
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。