arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过系统多智能体深度强化学习实现复杂环境下的多无人机协同导航

Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning

Yu Su, Nabil Aouf

arXiv 2607.25754首次发表:更新:

AI 中文总结

针对复杂环境下多无人机协同导航面临的问题,提出多智能体深度强化学习框架,通过协同探索等方法解决,含感知、示范缓冲区和课程调度机制,能抽象特征实现跨场景转移,经仿真验证有良好性能。

AI 中文摘要

复杂环境下多智能体无人机的协同导航面临局部最优陷阱、稀疏奖励、智能体间学习不平衡和跨场景泛化不足等关键挑战。本文提出了一个多智能体深度强化学习框架,通过协同探索、示范利用、安全课程调度和结构感知泛化来解决这些问题。该框架包含感知机制、分层协作示范缓冲区和安全感知双条件课程调度机制,还能通过抽象局部几何特征实现跨场景转移。在混合静态 - 动态障碍物设置下验证了其对动态干扰的强大适应性,仿真结果证实了其在协作成功率、导航鲁棒性、零样本跨场景泛化和动态环境适应性方面的强大性能。

英文摘要

Cooperative navigation of multi-agent UAVs in complex environments faces key challenges including local optima traps, sparse rewards, learning imbalance among agents, and insufficient cross-scenario generalisation. This paper proposes a multi-agent deep reinforcement learning framework that addresses these issues through coordinated exploration, demonstration exploitation, safe curriculum scheduling, and structure-aware generalisation. First, a perception mechanism combining memory of visited states, directional novelty estimates, and penalty backpropagation enables agents to proactively detect and escape local optima. Second, a hierarchical collaborative demonstration buffer with tiered behaviour cloning manages trajectories by degree of team collaboration and applies differential supervision to the actor network, improving demonstration utilisation under sparse collaborative signals. Third, a safety-aware dual-condition curriculum scheduling mechanism reviews mastered scenarios through back-testing and experience pre-filling during training, suppressing catastrophic forgetting while ensuring both task performance and flight safety. For generalisation, local geometric features computed from sensor readings are abstracted into a domain parameter, through which a structure-aware gating network and mixture-of-experts mechanism condition the policy on local structural patterns rather than scenario-specific coordinates, enabling cross-scenario transfer without exposure to the target environment. The framework is further validated under mixed static-dynamic obstacle settings, showing robust adaptability to dynamic disturbances. Simulation results confirm strong performance in collaboration success rate, navigation robustness, zero-shot cross-scenario generalisation, and dynamic environment adaptability.

Comments13 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑