arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向实用精准农业:嵌入式边缘硬件上的实时水果检测与视频分析

Towards Practical Precision Agriculture: Real-Time Fruit Detection and Video Analytics on Embedded Edge Hardware

Ivica Dimitrovski, Vlatko Spasev, Ivan Kitanovski, Petre Lameski, Dane Boshev

arXiv 2609.13551首次发表:更新:

发表机构

Faculty of Computer Science and Engineering, University Ss Cyril and Methodius; Faculty of Agricultural Sciences and Food, University Ss Cyril and Methodius(圣西里尔和迪乌斯大学计算机科学与工程学院; 圣西里尔和迪乌斯大学农业科学与食品学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出在Jetson Orin Nano Super上集成YOLO26s与DeepStream的实时水果检测跟踪计数框架,验证了FP16精度下高效能与低能耗,并指出采集几何对跟踪稳定性的关键影响。

AI 中文摘要

静态图像基准无法捕捉实际果园视频分析中的计算和时间需求。本研究提出了一个在NVIDIA Jetson Orin Nano Super上实现实时水果检测、跟踪和计数的端到端框架。一个轻量级的YOLO26s检测器在统一协议下,分别在代表苹果、芒果、蓝莓和草莓的四个公共数据集上独立训练。这些模型使用PyTorch和TensorRT以FP32、FP16和INT8精度部署在嵌入式平台上。随后使用APPLE MOTS进行时间视频分析,因为它提供了具有持久水果身份的果园序列,从而能够评估多目标跟踪和唯一水果计数。所选的FP16 TensorRT检测器被集成到NVIDIA DeepStream流水线中,该流水线结合了硬件加速解码、ByteTrack跟踪和基于运动感知的越线分析。在四个检测任务中,平均测试mAP@50:95的范围为0.4957至0.8656。在Jetson上,TensorRT FP16在13.41-14.98毫秒的预测延迟下实现了66.76-74.56帧/秒的速度,同时相对于PyTorch FP32,mAP@50:95仅降低了0.0020-0.0054,总能耗降低了约64-66%。完整的检测-跟踪-分析流水线达到了44.96-54.11 FPS,并在无输出帧丢失的情况下维持了配置的30 FPS输入速率。在留出的果园视频序列上,HOTA的范围为0.345至0.538,事件级计数F1为0.611至0.803,相对计数误差为6.2%至51.6%。性能因采集几何形状而异:近侧向行观测产生最稳定的跟踪和计数,而前向穿越尽管采用了空间自适应计数几何,仍然受到关联和召回率的限制。这些结果表明,实用的基于边缘的水果监测需要高效的检测和能够支持可靠时间关联的采集几何形状。

英文摘要

Static-image benchmarks do not capture the computational and temporal requirements of practical orchard video analytics. This study presents an end-to-end framework for real-time fruit detection, tracking, and counting on the NVIDIA Jetson Orin Nano Super. A lightweight YOLO26s detector is trained independently on four public datasets representing apples, mangoes, blueberries, and strawberries under a common protocol. The models are deployed on embedded platform using PyTorch and TensorRT at FP32, FP16, and INT8 precision. APPLE MOTS is then used for temporal video analytics because it provides orchard sequences with persistent fruit identities, enabling evaluation of multi-object tracking and unique-fruit counting. The selected FP16 TensorRT detector is integrated into an NVIDIA DeepStream pipeline combining hardware-accelerated decoding, ByteTrack tracking, and motion-aware line-crossing analytics. Across the four detection tasks, mean test mAP@50:95 ranges from 0.4957 to 0.8656. On the Jetson, TensorRT FP16 achieves 66.76-74.56 images/s at 13.41-14.98 ms prediction latency, while reducing mAP@50:95 by only 0.0020-0.0054 and gross energy consumption by approximately 64-66% relative to PyTorch FP32. The complete detector-tracker-analytics pipeline reaches 44.96-54.11 FPS and sustains the configured 30-FPS input rate without output-frame loss. On held-out orchard video sequences, HOTA ranges from 0.345 to 0.538, event-level counting F1 from 0.611 to 0.803, and relative count error from 6.2% to 51.6%. Performance varies across acquisition geometries: near-lateral row viewing yields the most stable tracking and counting, whereas forward traversal remains association- and recall-limited despite spatially adaptive counting geometry. These results show that practical edge-based fruit monitoring requires efficient detection and acquisition geometries that support reliable temporal association.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑