学习飞行:基于紧凑目标中心线索与强化学习的稳定视觉引导无人机伺服控制
Learning to Fly: Stable Vision-Guided UAV Servoing with Compact Target-Centric Cues and Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本研究提出利用紧凑目标中心线索与低维传感器测量,结合课程强化学习实现无人机长时程视觉伺服,在鲁棒性和抗扰动方面优于经典控制器。
中文摘要 AI 辅助
视觉引导的无人机强化学习因策略优化不稳定、探索激进以及高维视觉感知的成本而仍然具有挑战性。在本工作中,我们研究了利用紧凑目标中心线索结合低维传感器测量的长时程无人机视觉伺服控制。我们并非直接从RGB图像学习,而是通过轻量级目标分割提供图像空间偏移和相对深度,并将其与四旋翼速度和投影重力测量相结合,构成紧凑的12维策略观测。我们将直接PPO与三种匹配预算的课程策略进行比较:视觉课程逐步扩大目标放置难度,动力学课程逐步放宽动作约束和平滑,以及联合课程结合两种进展。所有策略在标称性能上相当,在跟踪指标上具有互补优势。观测消融实验表明,本体感觉测量对稳定飞行至关重要,图像空间线索对目标对齐至关重要,而在评估设置中,显式深度并非强性能所必需。与调优的经典视觉伺服控制器相比,学习策略在强控制和视觉扰动下表现出更强的鲁棒性,而视觉课程在未见目标运动下表现出最小的性能下降。总体而言,结果表明紧凑目标中心表示能够支持鲁棒的长时间空中视觉伺服,并且视觉课程训练能够在标称性能提升有限的情况下提高对动态分布偏移的鲁棒性。
英文摘要
Vision-guided reinforcement learning for Unmanned Aerial Vehicles (UAVs) remains challenging due to unstable policy optimisation, aggressive exploration, and the cost of high-dimensional visual perception. In this work, we investigate long-horizon UAV visual servoing using compact target-centric cues combined with low-dimensional sensor measurements. Rather than learning directly from RGB images, lightweight target segmentation provides image-space offsets and relative depth, which are combined with quadrotor velocity and projected-gravity measurements into a compact 12D policy observation. We compare Direct PPO with three matched-budget curriculum strategies: a Visual curriculum that progressively expands target placement difficulty, a Dynamics curriculum that gradually relaxes action constraints and smoothing, and a Joint curriculum that combines both progressions. All strategies reach comparable nominal performance, with complementary advantages across tracking metrics. Observation ablations show that proprioceptive measurements are critical for stable flight and image-space cues for target alignment, while explicit depth is not necessary for strong performance in the evaluated setting. Against tuned classical visual-servo controllers, learned policies show greater robustness to strong control and visual perturbations, while the Visual curriculum exhibits the smallest degradation under unseen target motion. Overall, the results demonstrate that compact target-centric representations can support robust long-horizon aerial visual servoing and that visual curriculum training can improve robustness to dynamic distribution shifts despite limited gains in nominal performance.
发表机构
- Indian Institute of Technology Bombay(印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。