发表机构
National University of Singapore; ST Engineering Unmanned & Integrated Systems Pte. Ltd.(新加坡国立大学; 胜科宇航无人及集成系统私人有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于物理的GPU仿真与结构化多智能体强化学习框架,训练去中心化MAPPO策略,实现拖船-驳船协同操纵,性能优于PID和集中式PPO,并支持零样本泛化与扩展。
AI 中文摘要
自主拖船作业对于自动化港口物流和船舶操纵等海事操作至关重要,其中多艘拖船必须协同运输/操纵一艘较大的船舶。在此场景中,协同推动因耦合水动力学、低阻力、强环境扰动、欠驱动驳船动力学以及接触丰富的交互而具有挑战性。传统控制方法通常依赖简化模型和固定配置,这限制了其适应性,而基于学习的方法则受限于缺乏可扩展且物理真实的训练环境。我们通过引入一个基于物理的、GPU加速的仿真与学习框架来解决这些挑战,用于协同拖船操纵。我们的仿真器包含自定义浮力模型、波浪建模和水动力阻力,并支持在海洋动力学下进行大规模多智能体训练。在该仿真器中,我们训练了一个去中心化的MAPPO(多智能体PPO)策略,并辅以结构化控制先验(SCP)以提高训练稳定性并保持可行的推动配置。我们在直线航行、转向和减速任务上评估了所学策略,结果表明,与基于PID的控制器和集中式PPO基线相比,我们的去中心化框架产生了更可靠、更准确的操纵性能。我们进一步展示了在更具挑战性的海况和高级机动中的零样本泛化能力,以及尽管仅使用两个智能体训练,仍能零样本扩展到三艘和四艘拖船的更大团队。
英文摘要
Autonomous tugboating is central for automating maritime operations such as port logistics and vessel maneuvering, where multiple tugboats must cooperatively transport/manipulate a larger vessel. Collaborative pushing in this setting is challenging due to coupled hydrodynamics, low resistance, strong environmental disturbances, underactuated barge dynamics, and contact-rich interactions. Conventional control methods often rely on simplified models and fixed configurations, which limit their adaptability, while learning-based approaches are constrained by the lack of scalable and physically realistic training environments. We address these challenges by introducing a physics-based, GPU-accelerated simulation and learning framework for collaborative tugboat manipulation. Our simulator incorporates a customized buoyancy model, wave modeling, and hydrodynamic resistance, and supports large-scale multi-agent training under marine dynamics. In this simulator, we train a decentralized MAPPO (Multi-Agent PPO) policy augmented with a structured control prior (SCP) to improve training stability and maintain feasible pushing configurations. We evaluate our learned policy on straight-line transit, turning, and deceleration tasks, where we show that our decentralized framework yields more reliable and accurate maneuvering performance compared to a PID-based controller and a centralized PPO baseline. We further demonstrate zero-shot generalization to more challenging sea states and advanced maneuvers, as well as zero-shot scalability to larger teams of three and four tugboats despite training with only two agents.
CommentsSubmitted to DAI