arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PAANI:面向河流机器人仿真的设备端视觉证据融合与可解释引导

PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation

Savio Cardoz, Santhiya Rajan

arXiv 2609.22353首次发表:更新:

发表机构

ACL Digital, Pune; Multiverse Computing, Spain; PSG College of Technology, Coimbatore, India(ACL数字,浦那; Multiverse Computing,西班牙; PSG技术学院,哥印拜陀,印度)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PAANI提出一种设备端视觉证据融合架构,结合YOLO11n检测器与MobileNetV3 Small分割器,在Arduino UNO Q上实现可解释的河流机器人引导,实验验证了模型精度与延迟性能。

AI 中文摘要

移动河流监测机器人必须解读地理航点无法描述的障碍物和水域边界。在资源受限的平台上,将不完美的视觉预测转换为及时且可检查的引导是一个独特的挑战。一个物体标签或转向命令并不能解释哪些证据支持决策,或者何时该证据不可靠。我们提出了PAANI,一种设备端感知到引导的架构,它结合了项目训练的YOLO11n检测器和定制的MobileNetV3 Small语义分割器,并在Arduino UNO Q上进行了时间戳对齐的证据融合。有界跟踪提供物体持久性,而显式走廊策略结合了表面标签、接受的检测、紧迫性和掩膜不确定性。每个最终建议都公开其贡献证据和策略原因。ROS 2接口将本地AI管道连接到独立的Gazebo船只、定位和控制测试平台。训练使用10,000张WaterScenes图像进行四类检测,以及1,127张MaSTr1325图像进行分割,包括198张分割验证图像。选定的FP32 ONNX模型占用14.817 MB。检测器检查点在0.5 IoU下的测试mAP为0.7388,而单独评估的矩形ONNX导出在0.5 IoU下达到验证mAP为0.7367。分割ONNX验证mIoU为0.9750。五分钟的UNO Q记录在配置的0.5 Hz节奏下产生了467.8 ms和580.3 ms的中位数和第95百分位管道延迟。评估还识别了黑色输入误分类和采样率不匹配,这阻止了诊断性表观运动估计器收集足够的证据。这些结果支持一个可检查且可重用的边缘机器人基础,同时清晰地区分模型准确性和板载执行与经过验证的水上碰撞避免。

英文摘要

Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained platforms, converting imperfect visual predictions into timely and inspectable guidance is a distinct challenge. An object label or steering command does not explain which evidence supports a decision or when that evidence is unreliable. We present PAANI, an on-device perception to guidance architecture that combines a project trained YOLO11n detector and a custom MobileNetV3 Small semantic segmenter with timestamp aligned evidence fusion on Arduino UNO Q. Bounded tracking supplies object persistence, while an explicit corridor policy combines surface labels, accepted detections, urgency and mask uncertainty. Each final advisory exposes its contributing evidence and policy reasons. ROS 2 interfaces connect the local AI pipeline to a separate Gazebo vessel, localization and control testbed. Training uses 10,000 WaterScenes images for four-class detection and 1,127 MaSTr1325 images for segmentation, including 198 segmentation validation images. The selected FP32 ONNX models occupy 14.817 MB. Detector checkpoint test mAP at 0.5 IoU is 0.7388, while the separately evaluated rectangular ONNX export achieves validation mAP at 0.5 IoU of 0.7367. Segmentation ONNX validation mIoU is 0.9750. A five-minute UNO Q recording produced median and 95th percentile pipeline latencies of 467.8 ms and 580.3 ms at a configured 0.5 Hz cadence. The evaluation also identifies black input misclassification and a sampling rate mismatch that prevents the diagnostic apparent motion estimator from collecting sufficient evidence. These results support an inspectable and reusable edge robotics foundation while clearly distinguishing model accuracy and on-board execution from validated on-water collision avoidance.

Comments20 pages, 8 figures, 11 tables. Includes system and AI architecture diagrams, model-training results, qualitative evaluations, and Arduino UNO Q deployment measurements. Project code, trained models, ONNX artifacts, logs, and reproducibility documentation are available at https://github.com/immanuelihs/ASV_PAANI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑