arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统一动力学框架:NVIDIA Isaac Sim中六自由度管道跟踪ROV的强化学习与经典控制

A Unified Dynamics Framework for Reinforcement Learning and Classical Control of a Six-DOF Pipeline-Tracking ROV in NVIDIA Isaac Sim

Cheng Siong Chin, M. Venkateshkumar, Jianhua Zhang

arXiv 2610.04949首次发表:更新:

发表机构

Newcastle University Singapore; Amrita Vishwa Vidyapeetham; Qingdao University of Technology(新加坡纽卡斯尔大学; 阿姆里塔大学; 青岛理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一个统一动力学框架,在NVIDIA Isaac Sim中训练六自由度ROV的管道跟踪策略,通过共享USD场景和Coriolis动力学模型,使强化学习与经典控制算法在相同条件下比较,实验表明PPO、TRPO和反馈线性化表现最佳。

AI 中文摘要

水下航行器的强化学习控制器通常针对一种物理表示进行训练,却在另一种物理表示下部署,因此报告的性能并不总能描述训练之外的行为。本文提出了一种用于六自由度遥控潜水器(ROV)的管道跟踪架构,在该架构中,一个通用场景描述(USD)场景为系统的两半部分提供真实的BlueROV2-Heavy质量、附加质量、阻尼、浮力和推进器参数:一个矢量化的NumPy实现的Fossen船舶方程,以及一个交互式的NVIDIA Isaac Sim部署,该部署在每一步都将相同的方程作为PhysX力应用。Coriolis-centripetal项是两个分支中的主要动力学模型;对PPO和TRPO进行的受控消融实验表明,包含该项不会使任一算法失稳,并适度改善了跟踪性能,TRPO的standoff RMS误差降低了约19%。五种强化学习算法,PPO、soft actor-critic、TD3、DDPG和TRPO,针对一个环境、奖励和随机化评估框架进行训练,通过一个检查点兼容层,用相同的代码对任何策略进行评分。该流程扩展了六个经典基线,PID、滑模、模糊逻辑、反馈线性化、模型预测控制和自适应神经模糊推理系统,它们由与学习策略相同的制导几何和推进器分配驱动。在启用Coriolis的动力学下,PPO、TRPO和反馈线性化达到了100%成功率和竞争性精度的最强组合;PID、模糊控制和神经模糊基线也达到了100%成功率,但跟踪较宽松;DDPG和TD3各自表现出一种特定的、可解释的失败模式,而非离线策略学习的一般弱点;经典控制仍然是针对最佳学习策略的强大基线。

英文摘要

Reinforcement learning controllers for underwater vehicles are usually trained against one physics representation and deployed against another, so reported performance does not always describe behavior outside training. This paper presents a pipeline-tracking architecture for a six-degree-of-freedom remotely operated vehicle (ROV) in which one Universal Scene Description (USD) scene supplies the real BlueROV2-Heavy mass, added-mass, damping, buoyancy, and thruster parameters to both halves of the system: a vectorized NumPy implementation of Fossen's marine-craft equations, and an interactive NVIDIA Isaac Sim deployment applying the identical equations as PhysX forces at every step. The Coriolis-centripetal term is the primary dynamics model in both branches; a controlled ablation on PPO and TRPO shows that including it does not destabilize either algorithm and modestly improves tracking, about 19 percent lower standoff RMS error for TRPO. Five reinforcement learning algorithms, PPO, soft actor-critic, TD3, DDPG, and TRPO, are trained against one environment, reward, and randomized evaluation harness through a checkpoint-compatibility layer scoring any policy with the same code. The pipeline is extended with six classical baselines, PID, sliding-mode, fuzzy logic, feedback linearization, model predictive control, and an adaptive neuro-fuzzy inference system, driven by the same guidance geometry and thruster allocation as the learned policies. Under Coriolis-enabled dynamics, PPO, TRPO, and feedback linearization reach the strongest combination of 100 percent success and competitive accuracy; PID, fuzzy control, and the neuro-fuzzy baseline also reach 100 percent success with looser tracking; DDPG and TD3 each show a specific, explainable failure mode rather than a general weakness of off-policy learning; and classical control remains a strong baseline against the best learned policies.

Comments26 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑