RALLY:面向智能体无人机集群的角色自适应大语言模型驱动共轭导航
RALLY: Role-Adaptive LLM-Driven Yoked Navigation for Agentic UAV Swarms
- Hangzhou Institute for Advanced Study, UCAS(杭州高等研究机构,中国科学院大学)
- Zhejiang University(浙江大学)
- Zhejiang Lab(浙江实验室)
- Macau University of Science and Technology(澳门科学理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对无人机集群传统控制方法泛化差、LLM控制框架探索不足的问题,提出角色自适应LLM驱动的RALLY算法,融合语义通信、动态角色机制与半离线训练,提升集群导航任务性能。
AI中文摘要:
无人机(UAV)集群的智能控制已成为关键研究热点,该领域通常要求集群在避障的同时实现有效导航,并对多个任务目标达成持续覆盖。传统多智能体强化学习(MARL)方法虽具备动态适应性,但受限于数值通信中的语义鸿沟以及同构角色结构的僵化问题,导致泛化性较差且任务可扩展性有限。近期基于大语言模型(LLM)的控制框架借助海量先验知识展现出强大的语义推理能力,但这类工作因缺乏在线学习且过度依赖静态先验,往往难以实现有效探索,进而造成个体潜力与整体系统性能下降。为解决上述局限,我们提出了角色自适应LLM驱动共轭导航算法RALLY。具体而言,我们首先构建了一个LLM驱动的语义决策框架,采用结构化自然语言实现高效语义通信与协作推理;随后引入动态角色异构机制,支持自适应角色切换与个性化决策;此外,我们提出了基于角色价值混合网络(RMIX)的分配策略,将LLM离线先验与MARL在线策略相融合,实现角色选择策略的半离线训练。在多智能体粒子环境(MPE)和软件在环(SITL)平台上的实验表明,RALLY在任务覆盖率、收敛速度以及泛化性方面均优于传统方法,凸显了其在多无人机智能体系统协作导航领域的巨大应用潜力。
英文摘要:
Intelligent control of Unmanned Aerial Vehicles (UAVs) swarms has emerged as a critical research focus, and it typically requires the swarm to navigate effectively while avoiding obstacles and achieving continuous coverage over multiple mission targets. Although traditional Multi-Agent Reinforcement Learning (MARL) approaches offer dynamic adaptability, they are hindered by the semantic gap in numerical communication and the rigidity of homogeneous role structures, resulting in poor generalization and limited task scalability. Recent advances in Large Language Model (LLM)-based control frameworks demonstrate strong semantic reasoning capabilities by leveraging extensive prior knowledge. However, due to the lack of online learning and over-reliance on static priors, these works often struggle with effective exploration, leading to reduced individual potential and overall system performance. To address these limitations, we propose a Role-Adaptive LLM-Driven Yoked navigation algorithm RALLY. Specifically, we first develop an LLM-driven semantic decision framework that uses structured natural language for efficient semantic communication and collaborative reasoning. Afterward, we introduce a dynamic role-heterogeneity mechanism for adaptive role switching and personalized decision-making. Furthermore, we propose a Role-value Mixing Network (RMIX)-based assignment strategy that integrates LLM offline priors with MARL online policies to enable semi-offline training of role selection strategies. Experiments in the Multi-Agent Particle Environment (MPE) environment and a Software-In-The-Loop (SITL) platform demonstrate that RALLY outperforms conventional approaches in terms of task coverage, convergence speed, and generalization, highlighting its strong potential for collaborative navigation in agentic multi-UAV systems.