RTK-Vision PPO 用于空中航母上微型无人机自主回收
RTK-Vision PPO for Autonomous Micro UAV Recovery on an Airborne Carrier
查看机构详情
- Indian Institute of Technology Hyderabad(印度理工学院海得拉巴分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出RTK视觉引导的PPO强化学习框架,实现微型无人机在移动空中航母上的自主起飞、出击与回收,仿真成功率99.55%,室外任务成功率92.9%。
中文摘要 AI 辅助
微型无人机(UAV)在移动空中航母上的自主回收,使得可重复使用的部署-任务-回收操作成为可能,但该过程耦合了远距离会合、近距离感知、航母运动、气动相互作用以及不连续的接触事件。本文提出了一种RTK视觉引导的强化学习框架,其中子无人机由较大的航母物理搭载,在航母飞行过程中从航母上起飞,执行独立出击任务,返回航母当前位置,重新对接,随后随航母一起下降。两架飞行器均携带RTK-GNSS,航母持续与子无人机共享其导航状态。在回收甲板附近,RTK保持激活状态,同时一个带有基准标记检测器的下视摄像头提供基于标记的相对对准提示。控制终端回收阶段的近端策略优化(PPO)策略,在基于物理的MuJoCo仿真环境中训练,该环境包含显式传感器噪声模型、气动干扰替代模型和标记延迟随机化,随后迁移至硬件。PX4保留低级稳定控制,一个确定性安全门独立于学习策略授权下降。PPO检查点在2000个留出的随机终端情节中达到99.55%的成功率,而相同条件下调优的PD基线为78.4%,中位平面终端误差为6.62厘米。在14次室外试验中,完整任务在13次试验中成功(92.9%),涵盖近区域回收以及航母从释放点平移后的回收。结果表明,这是一个完整的自主空中部署与回收循环,而非孤立的着陆机动,为检查、监视和移动物流应用中可重复使用的航母-子机操作奠定了实用基础。
英文摘要
Autonomous recovery of a micro unmanned aerial vehicle (UAV) onto a moving airborne carrier enables reusable deploy-mission-recover operation, but couples long-range rendezvous, close-range perception, carrier motion, aerodynamic interaction, and a discontinuous contact event. This paper presents an RTK-vision-guided reinforcement-learning framework in which a child UAV is physically transported by a larger carrier, takes off from the carrier while airborne, executes an independent sortie, returns to the carrier's current position, redocks, and subsequently descends with the carrier. Both vehicles carry RTK-GNSS, and the carrier continuously shares its navigation state with the child. Near the recovery deck, RTK remains active while a downward-facing camera with a fiducial marker detector provides marker-relative alignment cues. A proximal policy optimization (PPO) policy governing the terminal recovery phase is trained in a physics-based MuJoCo simulation environment with explicit sensor noise models, an aerodynamic disturbance surrogate, and marker-latency randomization, then transferred to hardware. PX4 retains low-level stabilization, and a deterministic safety gate authorizes descent independently of the learned policy. The PPO checkpoint achieves 99.55% success over 2,000 held-out randomized terminal episodes, compared with 78.4% for a tuned PD baseline under identical conditions, with a median planar terminal error of 6.62 cm. Across 14 outdoor trials, the full mission succeeds in 13 trials (92.9%), spanning both near-region recovery and recovery after the carrier translates away from the release point. The results demonstrate a complete autonomous aerial deployment-and-recovery cycle rather than an isolated landing maneuver, establishing a practical basis for reusable carrier-child operation in inspection, surveillance, and mobile-logistics applications.