arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20325cs.CV

AgriScope:面向农业图像的像素级多模态理解

AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

  • Khalifa University of Science and Technology(哈利法科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed

AI总结:

针对农业图像理解缺乏像素级定位的问题,提出统一框架AgriScope,结合生物语义编码与密集空间定位,并构建含50万图像、1100万样本的AgriGround数据集,在多项任务上验证了有效性。

AI中文摘要:

农业图像理解需要在复杂的现实条件下对植物病害、害虫、作物结构和植物物种进行细粒度识别。尽管多模态大语言模型(MLLMs)近期取得了进展,但现有模型仍仅限于文本输出,缺乏像素级视觉定位能力。在这项工作中,我们引入了AgriScope,一个用于农业图像理解的统一像素级多模态框架。AgriScope在统一框架内联合支持图像级、区域级和像素级理解,使得农业图像能够执行诸如定位描述生成、指代表达分割和多轮多模态交互等任务。AgriScope通过生物语义编码、密集空间表示和像素解码,将生物学专业语义表示与密集空间定位相结合。为了支持大规模定位学习,我们引入了AgriGround,一个大规模像素级农业多模态指令微调数据集,包含超过50万张图像和1100万条指令跟随样本,涵盖植物病害分析、作物与杂草识别、害虫识别和细粒度植物学理解。AgriGround通过一个多阶段自动标注流程构建,该流程整合了多模态描述生成、短语级定位、分割掩码生成和任务导向指令合成,以产生密集定位的监督信号。在多个农业视觉语言任务上的大量实验证明了AgriScope在像素级多模态理解方面的有效性,为农业视觉语言学习和视觉定位建立了强有力的基准。数据集和代码将在(此https URL)公开发布。

英文摘要:

Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Multimodal Large Language Models (MLLMs), existing models remain limited to text-only outputs and lack pixel-level visual grounding capabilities. In this work, we introduce AgriScope, a unified pixel-grounded multimodal framework for agricultural image understanding. AgriScope jointly supports image-level, region-level, and pixel-level understanding within a unified framework, enabling tasks such as grounded caption generation, referring expression segmentation, and multi-turn multimodal interaction for agricultural imagery. AgriScope integrates biologically specialized semantic representations with dense spatial grounding through biological-semantic encoding, dense spatial representations, and pixel decoding. To support large-scale grounded learning, we introduce AgriGround, a large-scale pixel-grounded agricultural multimodal instruction-tuning dataset containing over 500K images and 11M instruction-following samples spanning plant disease analysis, crop and weed identification, insect pest recognition, and fine-grained botanical understanding. AgriGround is constructed through a multi-stage automatic annotation pipeline that integrates multimodal caption generation, phrase-level grounding, segmentation mask generation, and task-oriented instruction synthesis to produce densely grounded supervision. Extensive experiments across multiple agricultural vision-language tasks demonstrate the effectiveness of AgriScope in pixel-grounded multimodal understanding, establishing a strong benchmark for agricultural vision-language learning and visual grounding. The dataset and code will be made publicly available at (https://github.com/boudiafA/AgriScope)

↑