arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14135cs.ROcs.LG

AgilePE:基于自博弈强化学习的自主无人机追逃

AgilePE: Autonomous UAV Pursuit-Evasion via Self-Play Reinforcement Learning

  • Zhejiang University(浙江大学)
  • Tsinghua University(清华大学)
  • Chongqing University(重庆大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenhao Tang, Tianyang Chen, Zhejun Cui, Boyuan An, Jiayu Chen, Ruize Zhang, Huidong Liu, Tianyue Wu, Qingmin Liao, Fei Gao, Yu Wang, Chao Yu

AI总结:

该研究提出AgilePE系统,通过自博弈强化学习实现自主无人机追逃,其策略可零样本迁移至真实四旋翼,展现出复杂追逃战术与交互式部署能力。

AI中文摘要:

自主追逃是无人机(UAV)面临的基础挑战,需在紧密耦合的动力学与不断变化的对手行为下实现快速决策。传统基于规则或微分博弈的方法往往难以应对高维空中交互与敏捷机动。本文提出AgilePE,一个基于自博弈强化学习的自主无人机追逃完整系统。AgilePE在统一框架中整合了敏捷底层控制、竞争策略优化与现实迁移部署。策略直接将机载状态观测映射至总推力与机体速率(CTBR)指令,实现端到端敏捷机动,无需中间轨迹规划器或航路点控制器。训练阶段采用带优先虚构自博弈(PFSP)的竞争自博弈及多样化对手池,使智能体可针对历史策略提升能力,同时稳定优化过程并减少策略振荡,最终涌现出复杂的追逃策略。现实部署时,开发了硬件对齐的仿真流水线,对执行器响应动力学、通信延迟与域随机化进行建模,学习到的策略可零样本迁移至真实四旋翼,无需任务特定调参。现实实验复现了仿真中观测到的追逃战术,包括快速规避与侧翼包抄,并展示了双智能体的交互式零样本部署。

英文摘要:

Autonomous pursuit-evasion is a fundamental challenge for Unmanned Aerial Vehicles (UAVs), requiring rapid decision-making under tightly coupled dynamics and continuously changing opponent behaviors. Traditional rule-based or differential-game approaches often struggle with high-dimensional aerial interactions and agile maneuvering. We present AgilePE, a complete system for autonomous UAV pursuit-evasion via self-play reinforcement learning. AgilePE integrates agile low-level control, competitive policy optimization, and sim-to-real deployment in a unified framework. The policy directly maps onboard state observations to Collective Thrust and Body Rates (CTBR) commands, enabling end-to-end agile maneuvering without intermediate trajectory planners or waypoint controllers. For training, we use competitive self-play with Prioritized Fictitious Self-Play (PFSP) and a diversified opponent pool, enabling agents to improve against historical policies while stabilizing optimization and reducing policy oscillation. This process leads to the emergence of sophisticated pursuit and evasion strategies. For real-world deployment, we develop a hardware-aligned simulation pipeline that models actuator-response dynamics, communication latency, and domain randomization. The learned policies transfer zero-shot to real quadrotors without task-specific tuning. Real-world experiments reproduce pursuit-evasion tactics observed in simulation, including rapid dodging and flanking, and demonstrate interactive two-agent zero-shot deployment.

补充信息

↑