可配置的多阶段视觉流水线用于作物病虫害诊断
Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis
浏览论文内容
中文总结 AI 辅助
针对现有作物病虫害诊断系统不可配置且性能不足的问题,提出三阶段可配置视觉流水线,用微调模型或小型专家模型实现,提升准确率并支持弃权(不执行)与请求重拍。
中文摘要 AI 辅助
this http URL 是 Digital Green 为小农户提供的农场咨询服务。当作物出现异常时,农民拍摄照片并发送,这张照片就是全部问题:没有症状描述,没有作物名称,通常甚至没有文字。该服务必须从在田间、光线不足且相机移动的情况下用廉价手机拍摄的图像中,判断照片是否可用、显示何种作物以及存在什么问题。当前执行此任务的系统无法调整。它没有用于照片拒绝的可调阈值,无法添加作物和问题,也没有可设置的置信度截断值。我们研究了来自埃塞俄比亚、印度、肯尼亚和尼日利亚发送至 this http URL 的约116万张照片。生产质量门拒绝了其判断图像的46.8%,超过四分之一的进入诊断的图像未返回作物名称,而在标记为“病害”的问题中,有35.8%是可无需作物即可识别的害虫。因此,我们将工作分为三个阶段:质量门(M0)、作物检测器(M1)和病害或害虫检测器(M2)。路线A使用一个微调的视觉语言模型(Qwen3-VL-4B)在单次调用中回答,填充所有三个阶段。路线B使用小型专家模型(DaViT、YOLO26)分别填充每个阶段。我们将生产中的GPT-4o质量门替换为小型MobileNetV3门,F1分数为86.9%,耗时12毫秒。在一个以相同方式对所有系统评分的测试集上,层次化DaViT-Base的作物准确率达到95.41%,而生产基线的准确率为91.46%。它在诊断方面也领先,且从不拒绝回答,而比较中的每个语言模型都留下了大量无诊断的行。微调模型保留了专家模型不具备的两项能力:一次调用完成所有三个阶段,以及在图像无法支持答案时请求更好的照片。
英文摘要
FarmerChat is Digital Green's farm advisory service for smallholder farmers. When something looks wrong with a crop, the farmer takes a photograph and sends it, and that photograph is the whole question: no symptom described, no crop named, often no text at all. The service has to determine whether the picture can be used, what crop it shows, and what is wrong with it, from images taken on cheap phones in a field, in poor light and with a moving camera. The system doing this today cannot be adjusted. It has no adjustable thresholds for photograph rejection, crops and problems cannot be added, and there is no confidence cut-off to set. We study about 1.16 million photographs sent to FarmerChat from Ethiopia, India, Kenya and Nigeria. The production quality gate rejected 46.8% of the images it judged, over a quarter of those reaching diagnosis returned no crop name, and 35.8% of the labelled problems filed under "disease" are pests, identifiable without the crop. We therefore split the work into three stages: a quality gate (M0), a crop detector (M1), and a disease or pest detector (M2). Route A fills all three with one fine-tuned vision-language model (Qwen3-VL-4B) answering in a single call. Route B fills each with a small specialist model (DaViT, YOLO26). We replace our production GPT-4o quality gate with a small MobileNetV3 gate at 86.9% F1 in 12 ms. On one test set scored the same way for every system, a hierarchical DaViT-Base achieves 95.41% crop accuracy against 91.46% for the production baseline. It also leads on diagnosis and never declines to answer, while every language model in the comparison leaves a large share of rows with no diagnosis. The fine-tuned model retains two capabilities the specialists do not have: one call for all three stages, and a request for a better photograph when the image cannot support an answer.
发表机构
- Digital Green(数字绿色)
机构由 AI 辅助整理,请以论文原文为准。