发表机构
MIT–WHOI Joint Program; Woods Hole Oceanographic Institution; Massachusetts Institute of Technology(麻省理工学院 - 伍兹霍尔海洋研究所联合项目; 伍兹霍尔海洋研究所; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究自主水下航行器细粒度控制与定位问题,提出利用CFD模型代理近似在强化学习中快速推理的方法,首次在6自由度AUV上成功部署零射击RL策略,相比传统控制器有能耗降低、速度加快、误差减小等优势。
AI 中文摘要
自主水下航行器(AUV)的细粒度控制和定位对于采样、维护及勘测应用至关重要。传统控制方法劳动强度大且对车辆配置或环境条件变化不稳健。强化学习(RL)有望快速开发控制器,通过域随机化(DR)处理一系列部署参数,但DR受基础模拟对真实物理建模能力限制。计算流体动力学(CFD)提供高保真阻力模型,但因计算开销难以在强化学习框架中利用。本文利用训练给定车辆CFD模型代理近似的想法,在RL管道内实现快速推理。首次在6自由度AUV上成功部署零射击RL策略,在基于CFD数据训练的代理阻力模型(SDM)上进行策略训练。与使用简化物理的控制器相比,能量使用降低31%,航点间移动速度快11%,误差减少19%。基于SDM的RL控制器更好地预测零射击转移,在奖励塑造设计选择上更稳健。使用DR完成参数受扰任务时,CFD策略是唯一成功转移的控制器。策略在受控水槽环境和实地进行评估,对其能力进行广泛测试。
英文摘要
Precise control and positioning of autonomous underwater vehicles (AUVs) is critical for sampling, maintenance, and survey applications. Reinforcement learning (RL) offers rapid controller development while handling a range of deployment parameters via domain randomization (DR). However, DR is limited by the underlying simulation's capacity to model real physics such as drag, which is a large contributor to sim-to-real gaps. Computational fluid dynamics (CFD) provides high-fidelity drag models but has significant computational cost. Thus, in this paper we train surrogate approximations of CFD data of a given vehicle, providing drag estimates 700,000x faster than solving for drag at each time step with CFD within the RL training pipeline. We deploy a Proximal Policy Optimization (PPO) policy zero-shot on a 6-DOF AUV in which policy training is performed on these surrogate drag models (SDMs), providing analysis on zero-shot transfer, reward shaping sensitivity, and physical parameter perturbations. On 15 identical segments in the ocean, we observe median gains of 19% in mean tracking error and 31% in control effort compared to the equivalent inertia box model. Our SDM based RL controller better predicts zero-shot transfer and reaches the most waypoints across all three reward shaping choices that we evaluate. Finally, we add 0.907 kg to the vehicle perturbing the mass, center-of-mass, and center-of-buoyancy. The policy reaches 100% of waypoints only when the SDM drag model is used; the policy fails to reach any of the 15 waypoints when the other drag models are used.
Comments16 pages, 13 figures