arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于起伏船甲板自主着陆的屏障形循环强化学习

Barrier-Shaped Recurrent Reinforcement Learning for Autonomous Landing on a Heaving Ship Deck

Ritwik Shankar, Chiranjeev Prachand, Abhishek, Soumya Ranjan Sahoo

arXiv 2610.03924首次发表:更新:

发表机构

Indian Institute of Technology Kanpur(印度理工学院坎普尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于循环策略的非对称actor-critic强化学习方法,结合控制障碍函数停止裕度,实现无人机在起伏船甲板上的自主着陆,并通过模拟和真实试验验证了高成功率与低接触速度。

AI 中文摘要

本文研究了无人航空器(UAV)旋翼机在起伏船甲板上的自主着陆问题,采用通过非对称actor-critic强化学习训练的循环策略:在训练期间,critic可观察未来6秒的甲板高度,而actor仅能获取机载传感器在飞行中提供的信息,即一个由飞行器相对于甲板的位置、速度和姿态组成的17维状态,并输出世界坐标系下的速度和偏航角速率指令。训练在4096个并行模拟环境中进行,随后在16个环境中进行微调,并集成机载视觉管线。在甲板相对垂直状态上引入了控制障碍函数(CBF)停止裕度,并在两个阶段应用:在训练期间作为奖励项,与未使用该策略训练的策略相比,它将模拟接触速度的中位数减半;在部署时作为运行时安全过滤器,在每个控制步骤进行评估,当裕度被违反时中止并重试下降。由于自动驾驶仪的解除武装逻辑无法检测飞行器停在移动甲板上,因此在着陆管线中直接命令在触地时进行基于接近度的推力切断。所提出的方法使用基于并联机械臂平台的甲板模拟器进行验证,该模拟器再现了船甲板的起伏运动(缩放至0.70米峰峰值,平均周期7.5秒),并使用带有板载摄像头的四旋翼无人机。在31次运动捕捉和20次基于视觉的试验中,无人机每次都成功着陆,中位接触时间分别为6.5秒和7.3秒;77%和50%的着陆在第一次尝试中成功,在推力切断瞬间(即策略控制下的最后速度)的平均甲板相对速度分别为0.29米/秒和0.27米/秒。补充视频:此https URL

英文摘要

This paper addresses autonomous landing of unmanned aerial vehicle (UAV) rotorcrafts on a heaving ship deck using a recurrent policy trained via asymmetric actor-critic reinforcement learning: the critic sees 6 s of future deck height during training, while the actor sees only what the onboard sensors provide in flight, a 17-dimensional state made of the vehicle's position, velocity and attitude relative to the deck and outputs world-frame velocity and yaw-rate commands. Training is performed across 4096 parallel simulated environments, followed by fine-tuning in 16 environments with an onboard vision pipeline in the loop. A control-barrier-function (CBF) stopping margin on the deck-relative vertical state is incorporated at two stages: as a reward term during training, where it halves the median simulated contact speed relative to a policy trained without it, and as a runtime safety filter at deployment, evaluated at every control step to abort and retry the descent when the margin is violated. Because the autopilot's disarm logic cannot detect the vehicle resting on a moving deck, proximity-based thrust cutoff at touchdown is commanded directly in the landing pipeline. The proposed approach is validated using a parallel-manipulator-platform-based deck emulator that reproduces the heaving motion of the ship deck (scaled to 0.70 m peak-to-peak, 7.5 s mean period) and a quadcopter UAV with an onboard camera. Across 31 motion-capture and 20 vision-based trials, the UAV landed every time, with median times to contact of 6.5 and 7.3 s; 77% and 50% landed on the first attempt, with mean deck-relative speeds of 0.29 and 0.27 m/s, respectively, at the instant of thrust cutoff, which is the last speed under the policy's control. Supplementary video: https://youtu.be/S5hDkrSZJt4

Comments8 pages, 8 figures, 1 table. Submitted to the 2027 IEEE International Conference on Robotics and Automation (ICRA). Video: https://youtu.be/S5hDkrSZJt4

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑