发表机构
Northwestern Polytechnical University; Wuhan Second Ship Design and Research Institute(西北工业大学; 武汉第二船舶设计研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对遥控水下机器人洋流扰动下的6自由度推力控制问题,提出TSRCA-PPO方法,经知识蒸馏学习后,在多性能指标上显著优于传统P-PID控制器。
AI 中文摘要
随着计算能力的持续提升,端到端强化学习在遥控水下机器人(ROV)控制领域得到了快速发展。然而,现有基于端到端强化学习的方法在洋流扰动下仍难以实现最优控制,尤其缺乏能同时满足低稳态跟踪误差、快速瞬态响应、节能运行及平滑推力输出的统一控制框架。为解决该问题,本文提出推力平滑-洋流快速自适应近端策略优化(TSRCA-PPO)方法,通过两阶段蒸馏学习框架学习近似最优策略,核心创新在于奖励函数设计与特权多编码器架构。 ablation研究验证了各模块的有效性,仿真结果表明,所提TSRCA-PPO方法在所有评估指标上均显著优于传统级联P-PID控制器:TSRCA-PPO将稳态位置误差、稳态姿态误差、调节时间、能量指标及推力平滑度指标分别降至对应P-PID值的42.7%、76.5%、10.6%、93.5%和15.9%。
英文摘要
With the continuous improvement of computational capabilities, end-to-end reinforcement learning has been rapidly developed for remotely operated vehicles control. Nevertheless, existing end-to-end reinforcement-learningbased methods still face challenges in achieving optimal control under oceancurrent disturbances. In particular, there remains a lack of a unified control framework that can simultaneously achieve low steady-state tracking error, rapid transient response, energy-efficient operation, and smooth controlforce outputs under disturbances. To address the issue, this paper proposes the thrust smoothness rapid current adaptation proximal policy optimization (TSRCA-PPO) method which learns a near-optimal strategy by a twostage distillation learning framework. The core innovations of this work lie in the reward-function design and the privileged multi-encoder architecture. Ablation studies validate the effectiveness of each module. Simulation results demonstrate that the proposed TSRCA-PPO method consistently outperforms the conventional cascaded P-PID controller across all evaluation metrics. Specifically, TSRCA-PPO reduces the steady-state position error, steady-state attitude error, settling time, energy index, and thrustsmoothness index to 42.7%, 76.5%, 10.6%, 93.5%, and 15.9% of the corresponding P-PID values, respectively.