OGR-MARL:面向受限港口水道中异构无人水面艇协同追踪的选项引导式残差多智能体强化学习
OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways
浏览论文内容
中文总结 AI 辅助
本文提出OGR-MARL框架,将其实例化为多款连续控制MARL算法,经厦门港水道实验验证,OGR-MASAC捕获率达75.0%,规则依从性与异构协同表现最优,且具备良好泛化潜力。
中文摘要 AI 辅助
受限港口水道中的异构无人水面艇(USV)协同追踪需要在导航、交通及角色约束下完成对逃逸目标的拦截。本文提出OGR-MARL,一种选项引导式残差多智能体强化学习框架,该框架与特定多智能体强化学习(MARL)算法解耦。OGR-MARL整合了共享逃逸目标信念、角色条件选项目标、自适应规则惩罚以及残差策略学习,使不同MARL算法可在规则引导行为的基础上学习校正动作,而非从头探索受限港口环境。本文以代表性连续控制MARL骨干算法实例化OGR-MARL,包括MADDPG、MATD3、MAPPO和MASAC,得到OGR-MADDPG、OGR-MATD3、OGR-MAPPO和OGR-MASAC。在抽象的厦门港水道场景实验中,OGR-MASAC实例化版本达到75.0%的捕获率,具备良好的任务有效性与规则依从性,且在测试方法中实现最优异构协同;无需重新训练,其零样本迁移至基于QGIS/AIS的厦门港地图即取得良好结果,证明OGR-MARL在更复杂港口场景中的泛化潜力。
英文摘要
Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a specific MARL algorithm. OGR-MARL integrates shared evader belief, role-conditioned option targets, adaptive rule penalties, and residual policy learning, allowing different MARL algorithms to learn corrective actions on top of rule-guided behaviors rather than exploring constrained port environments from scratch. We instantiate OGR-MARL with representative continuous-control MARL backbones, including MADDPG, MATD3, MAPPO, and MASAC, yielding OGR-MADDPG, OGR-MATD3, OGR-MAPPO, and OGR-MASAC. Experiments in an abstract Xiazhimen port-waterway scenario show that the OGR-MASAC instantiation achieves a 75.0% capture rate, promising mission-effective rule compliance, and the best heterogeneous coordination among the tested methods. Without retraining, zero-shot transfer to a QGIS/AIS-informed Xiazhimen map achieves promising results, demonstrating the generalization potential of OGR-MARL in more complex port scenarios.