arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向商业果园早期解剖结构青果分类的轻量级多模态视觉-语言框架

A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

Ranjan Sapkota, William Bu, Chen Chen, Yunjun Xu, Manoj Karkee

arXiv 2608.24935首次发表:更新:

发表机构

Cornell University; University of Central Florida; Institute of Artificial Intelligence (IAI)(康奈尔大学; 中佛罗里达大学; 人工智能研究所(IAI))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出适配TinyCLIP的轻量级多模态视觉-语言框架,实现复杂果园环境下幼果解剖结构细粒度分类,在NVIDIA T4上取得0.93宏F1,经优化可边缘部署,支撑自动化幼果分析与机器人疏果系统。

AI 中文摘要

准确识别苹果幼果的早期解剖结构(包括花萼、幼果本体和果柄),对机器人疏果、作物负载管理及其他精准果园作业至关重要。本研究提出一种轻量级多模态视觉-语言框架,该框架适配TinyCLIP以在复杂果园环境中实现幼果解剖结构的细粒度分类。从Scilate和Scifresh苹果园采集的600张高分辨率RGB图像构成数据集,被转换为224×224的图像块并标注为三个解剖类别。采用“一张某类的照片”这类领域特定语言提示,引导果园图像与园艺结构之间的多模态对齐。步长为112像素的滑动窗口推理策略将图像块级预测聚合为空间热图,实现与机器人疏果相关的幼果组件的可解释全图定位。在NVIDIA T4 GPU上进行的图像块级评估显示,花萼的F1分数为0.95,幼果本体为0.98,果柄为0.85,宏F1分数为0.93。使用ONNX和TensorRT进行面向部署的优化,使其能在NVIDIA Jetson硬件上高效推理,在INT8量化下保持精度,模型大小约为127-137 MB,图像块推理达到毫秒级。这些结果表明,轻量级视觉-语言模型可为自动化幼果分析及未来机器人疏果系统提供可解释且可边缘部署的感知能力,源代码和实现细节可在该httpsURL获取。

英文摘要

Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is essential for robotic thinning, crop-load management, and other precision orchard operations. This study presents a lightweight multimodal vision-language framework that adapts TinyCLIP for fine-grained fruitlet anatomy classification in complex orchard environments. A dataset of 600 high-resolution RGB images collected from Scilate and Scifresh apple orchards was converted into 224 x 224 image patches and annotated for three anatomical classes. Domain-specific language prompts, such as ``a photo of a class,'' were used to guide multimodal alignment between orchard imagery and horticultural structures. A sliding-window inference strategy with a stride of 112 pixels aggregates patch-level predictions into spatial heatmaps, enabling interpretable whole-image localization of fruitlet components relevant to robotic thinning. Patch-level evaluation on an NVIDIA T4 GPU achieved F1-scores of 0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle, with a macro-F1 score of 0.93. Deployment-oriented optimization using ONNX and TensorRT enabled efficient inference on NVIDIA Jetson hardware, preserved accuracy under INT8 quantization, and supported model sizes of approximately 127-137 MB with millisecond-level patch inference. These results demonstrate that lightweight vision-language models can provide interpretable and edge-deployable perception for automated fruitlet analysis and future robotic thinning systems. The source code and implementation details are publicly available at https://github.com/WilliamBu1/A-Lightweight-Vision-Language-Model-for-Early-Stage-Fruitlet-Classification-in-Apple-Orchards.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑