发表机构
North Carolina A&T State University(北卡罗来纳农工州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对无人机视觉制导的偏航角估计难题,提出基于YOLO边界框特征的可解释模糊推理框架,经实验验证其精度高、数据效率高且计算轻量,适用于实时无人机制导。
AI 中文摘要
无人机(UAV)基于视觉的制导至无人地面车辆(UGV)可支撑空-地协作机器人技术,但受传感不确定性、计算资源有限及控制可解释性需求影响,从机载视觉实现可靠连续偏航角估计仍具挑战性。现有深度学习与几何重建方法通常需要大量数据集、外部定位或复杂建模假设,降低了透明度且难以在资源受限平台部署。本文提出一种可解释模糊推理框架,从YOLO边界框提取的低维特征(目标质心位置、面积、长宽比)生成连续偏航角指令,无需显式几何建模。采用肩-三角-肩输入划分的Mamdani模糊系统作为可解释基线,其后为一阶Takagi-Sugeno模型,每个输入含三个前件隶属项,其参数由训练集分位数推导,形成紧凑的27条规则结构。实验使用VICON运动捕捉环境中的6169个标注样本,在5次随机训练-测试划分中,Takagi-Sugeno模型测试集平均绝对误差为0.140°±0.003°,均方根误差为0.200°±0.008°,最大绝对误差为1.254°±0.121°;±1°阈值内准确率为99.676%±0.270%,±3°与±5°阈值内准确率均为100.000%±0.000%;图像平面水平位移与预测偏航角符号的方向一致性达90.254%±0.612%。结果表明该框架透明、数据高效、计算轻量,适用于基于视觉的无人机对移动地面目标的实时制导。
英文摘要
Vision-based guidance of unmanned aerial vehicles (UAVs) toward unmanned ground vehicles (UGVs) supports cooperative aerial--ground robotics, but reliable continuous yaw estimation from onboard vision remains challenging because of sensing uncertainty, limited computation, and the need for interpretable control. Existing deep-learning and geometric-reconstruction approaches often require large datasets, external localization, or complex modeling assumptions, reducing transparency and deployment suitability on resource-constrained platforms. We present an interpretable fuzzy-inference framework that generates continuous yaw commands from low-dimensional features extracted from YOLO boxes: target centroid location, area, and aspect ratio. No explicit geometric modeling is required. A Mamdani fuzzy system serves as an interpretable baseline using a shoulder--triangle--shoulder input partition. It is followed by a first-order Takagi--Sugeno model with three antecedent membership terms per input, whose parameters are derived from training-set quantiles, yielding a compact 27-rule structure. Evaluation uses 6{,}169 labeled samples from a VICON motion-capture environment. Across five randomized train--test splits, the Takagi--Sugeno model achieves a test-set mean absolute error of $0.140^\circ \pm 0.003^\circ$, a root mean squared error of $0.200^\circ \pm 0.008^\circ$, and a maximum absolute error of $1.254^\circ \pm 0.121^\circ$. Within-threshold accuracies are $99.676% \pm 0.270%$ for $\pm1^\circ$ and $100.000% \pm 0.000%$ for both $\pm3^\circ$ and $\pm5^\circ$. Directional consistency between image-plane horizontal displacement and predicted yaw sign reaches $90.254% \pm 0.612%$. These results show that the framework is transparent, data-efficient, computationally lightweight, and suitable for real-time vision-based UAV guidance toward mobile ground targets.