基于视觉Transformer的非合作无人机合成到真实姿态估计
Synthetic-to-Real ViT-Based Pose Estimation of a Noncooperative UAV
浏览论文内容
中文总结 AI 辅助
本研究提出一种基于ViT(DINOv2骨干)的单目非合作无人机姿态估计方法,仅在合成数据上训练,通过姿态模糊感知策略和α-β滤波,在真实数据集上实现滤波后MAE降至8.74°,推理时间13.42毫秒。
中文摘要 AI 辅助
从图像中远程估计非合作无人机(UAV)的姿态至关重要,因为这些无人机无法提前被影响或安装仪器。基于深度学习的方法提供了有前景的解决方案;然而,其开发受到获取带有精确姿态标签的大规模真实世界数据集所需成本和难度的限制。合成图像提供了一种替代方案,但在合成数据上训练的模型必须克服合成到真实的域差距,才能泛化到真实世界图像。本研究探讨了基于视觉Transformer(ViT)模型的单目无人机姿态估计的固有合成到真实泛化能力。所提出的方法采用自监督的DINOv2骨干网络,并仅在标记的合成图像上训练,同时在标记的真实世界图像上评估。在训练和推理过程中引入了姿态模糊感知策略,以解决由三维目标投影到二维图像平面以及目标对称性引起的模糊性。推理阶段进一步集成了α-β滤波器以改进姿态估计。为了在操作要求下评估模型,使用包含77,077张标记无人机图像的真实数据集,在滤波前后分别评估平均角度误差(MAE)和推理时间。滤波前,模型达到的MAE为19.18°,推理时间为13.25毫秒;滤波后,这些值分别为8.74°和13.42毫秒。
英文摘要
Remote pose estimation of noncooperative Unmanned Aerial Vehicles (UAVs) from imagery is critical, as they cannot be influenced or instrumented in advance. Deep-learning-based approaches offer a promising solution; however, their development is constrained by the cost and difficulty of acquiring large-scale real-world datasets with accurate pose labels. Synthetic imagery provides an alternative, but models trained on synthetic data must overcome the synthetic-to-real domain gap to generalize to real-world imagery. This work investigates the inherent synthetic-to-real generalization capability of a Vision Transformer (ViT)-based model for monocular UAV pose estimation. The proposed approach employs a self-supervised DINOv2 backbone and is trained exclusively on labeled synthetic imagery while being evaluated on labeled real-world imagery. Pose ambiguity-aware strategies are incorporated during training and inference to address ambiguities arising from the projection of a three-dimensional target onto a two-dimensional image plane and from target symmetries. An $α$-$β$ filter is further integrated during inference to improve pose estimations. To assess the model under operational requirements, it is evaluated in terms of Mean Angular Error (MAE) and inference time, both before and after filtering, using a real-world dataset containing 77,077 labeled UAV images. Before filtering, the model achieves an MAE of $19.18^{\circ}$ and an inference time of $13.25$ ms, whereas after filtering, these values are $8.74^{\circ}$ and $13.42$ ms, respectively.
发表机构
- Naval Postgraduate School(海军研究生院)
机构由 AI 辅助整理,请以论文原文为准。