发表机构
Maastricht University(马斯特里赫特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对海上四旋翼无人机回收难题,提出基于课程学习和对抗风智能体的HAPPO强化学习方法,在分布外海况下显著提升成功率并降低坠毁率。
AI 中文摘要
在海上环境中回收无人机(UAV)因风湍流和船甲板运动而具有挑战性,使得传统降落方法往往变得不可靠,因此成为替代控制和学习方法的有价值测试案例。我们研究了由船载机械臂对四旋翼无人机进行模拟空中捕获,使用异构智能体近端策略优化(HAPPO)强化学习来学习鲁棒的协同控制策略。我们在NVIDIA Isaac Lab中使用课程学习和对抗风智能体(HARL-AC)进行HAPPO训练,并将获得的控制策略与基于课程域随机化生成的策略以及在单一海况下训练的基准进行比较。在海况0/4/5的分布内评估中,HARL-AC和域随机化的成功率相当,最高可达97.5%。在分布外的海况7/8/10中,HARL-AC泛化更好,在海况10下中位成功率最高提高16%,且与域随机化策略相比,坠毁率大幅降低,最高降低14%。此外,我们表明对抗训练的策略表现出更谨慎的行为,超时略微增加<3%,但在严重的未见条件下产生更安全的回收行为。
英文摘要
Recovering unmanned aerial vehicles (UAVs) in maritime environments is challenging due to wind turbulence and ship-deck motion, making it a valuable test case for alternative control and learning approaches as conventional landing approaches often become unreliable. We study simulated mid-air capture of quadrotor UAVs by a ship-mounted robotic arm, learning robust cooperative control policies with Heterogeneous-Agent Proximal Policy Optimization (HAPPO) Reinforcement Learning. We train with HAPPO using a curriculum and an adversarial wind agent (HARL-AC) in NVIDIA Isaac Lab, and compare the obtained control policies against those generated through curriculum-based domain randomization and a benchmark trained on a single sea state. In-distribution evaluation on sea states $0/4/5$ shows comparable success for HARL-AC and domain randomization of up to $97.5\%$. On out-of-distribution sea states $7/8/10$, HARL-AC generalizes better, achieving up to $16\%$ higher median success rate at sea state 10, and substantially lower crash rates of up to $14\%$ compared to the domain randomization policy. Furthermore, we show that the adversarially trained policy shows more cautious behavior, slightly increasing timeouts by $<3\%$, but yields safer recovery behavior in severe, unseen conditions.
Comments20 pages, 5 figures, submitted to BNAIC 2026