arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DRG-MAPPO:用于协同空战的分层动态角色图多智能体强化学习

DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat

Junlin Liu, Chengwei Li, Yang Gao, Hui Chang, Xinchen Zhang, Zhijun Zhao, Hao Zhao

arXiv 2609.11155首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, CASIA; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院自动化研究所; 中国科学院自动化研究所复杂系统认知与决策智能重点实验室; 中国科学院大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DRG-MAPPO框架,结合图关系建模与动态角色分配,实现协同空战中的战术协同,达到87%的胜率。

AI 中文摘要

多智能体强化学习(MARL)已成为自主系统和空战中复杂决策的关键范式。尽管MARL在空战中展现出巨大潜力,但实现复杂的战术协同仍是一项艰巨的挑战。这一困难主要归因于两个主要限制:(1)缺乏结构化的关系建模,导致智能体无法捕捉战场实体间复杂、时变的交互;(2)传统的扁平架构往往缺乏显式建模战术角色的能力,导致在高动态环境中任务分配模糊。为应对这些挑战,我们提出了分层动态角色图多智能体近端策略优化(DRG-MAPPO),一种新颖的MARL框架,将基于图的关系建模与动态角色分配相结合。具体而言,DRG-MAPPO构建了战场交互的图表示,并利用图注意力机制提取友方、敌方和威胁之间的关键关系特征。随后,高层策略采用动态角色分配机制来确定战术职责(如“领队”和“支援者”)。基于这些角色和编码的图关系特征,低层策略执行离散机动动作,促进战术策略与协同执行的联合优化。此外,设计了一个目标优先级辅助任务以促进集中火力等行为的涌现。实验结果表明,DRG-MAPPO达到了87%的最先进胜率,表明我们的框架有效平衡了协同空战中的关系建模、可解释性和优化稳定性。

英文摘要

Multi-Agent Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous systems and air combat. While MARL has demonstrated significant potential in air combat, achieving sophisticated tactical coordination remains a non-trivial challenge. This difficulty is largely attributed to two primary limitations: (1) the absence of structured relational modeling hinders agents from capturing complex, time-varying interactions among battlefield entities; and (2) conventional flat architectures often lack the capability to explicitly model tactical roles, leading to ambiguous task allocation in highly dynamic environments. To address these challenges, we propose Hierarchical Dynamic Role-Graph Multi-Agent Proximal Policy Optimization (DRG-MAPPO), a novel MARL framework that integrates graph-based relational modeling with dynamic role assignment. Specifically, DRG-MAPPO constructs a graph-based representation of battlefield interactions and leverages graph attention mechanisms to extract critical relational features among allies, enemies, and threats. Subsequently, a high-level policy employs a dynamic role assignment mechanism to determine tactical responsibilities (e.g., ``leader'' and ``supporter''). Conditioned on these roles and encoded graph-relational features, a low-level policy executes discrete maneuver actions, facilitating the joint optimization of tactical strategy and collaborative execution. Furthermore, a target-priority auxiliary task is designed to foster the emergence of behaviors such as focus-fire. Experimental results demonstrate that DRG-MAPPO achieves a state-of-the-art win rate of 87%, suggesting that our framework effectively balances relational modeling, interpretability, and optimization stability for cooperative air combat.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑