arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16326cs.MA

KC-BFPRL:基于知识引导的多无人机草原修复协作方法——通过 bilevel Formerpointer 强化学习实现

KC-BFPRL: Knowledge-Guided Multi-UAV Collaboration for Grassland Restoration via Bilevel Formerpointer-Based Reinforcement Learning

Dongbin Jiao, Xianyi Wang, Yuchen Yuan, Weibo Yang, Peng Yang, Peng Zhao, Zhanhuan Shang, Shi Yan

首次发表
浏览论文内容

中文总结 AI 辅助

KC-BFPRL 是一种知识引导的 bilevel Formerpointer 强化学习框架,可解决多无人机草原修复的 RAMP 问题,性能优于现有方法,适用于大规模自动化生态修复。

中文摘要 AI 辅助

多无人机(UAV)系统为草原生态系统修复等大规模环境任务提供了可扩展的服务平台。然而,协调机群作业需解决修复区域最大化问题(RAMP),这一非线性组合优化挑战因依赖载荷的能量动态特性及异质生态退化问题而变得复杂。我们提出一种新型知识引导的协作 bilevel Formerpointer 强化学习框架(KC-BFPRL)以应对该复杂性。KC-BFPRL 采用分层范式将 RAMP 分解为全局任务分配与局部修复规划,其中局部修复规划进一步分为上层轨迹规划与下层修复区域分配。其专用架构包含基于 Transformer 的编码器,用于融合静态环境特征与动态无人机状态,以及通过鲁棒演员-评论家框架训练的 Pointer Network 解码器。通过嵌入生态优先级规则与启发式逻辑,KC-BFPRL 实现了结构化热启动,解决了强化学习冷启动问题,同时确保严格满足约束条件。大量实验表明,KC-BFPRL 始终优于当前最优基线,取得了更优的目标值与效率;在最复杂场景 U8-R160 中,其保持了 0.00% 的最优性间隙,且运行速度约为 MAPDP 的三倍,验证了其在大规模自动化生态修复中的鲁棒性、可扩展性与实时适用性。

英文摘要

Multi-unmanned aerial vehicle (UAV) systems provide scalable service platforms for large-scale environmental tasks, such as grassland ecosystem restoration. However, coordinating fleet operations requires solving the restoration area maximization problem (RAMP). This non-linear combinatorial optimization challenge is complicated by payload-dependent energy dynamics and heterogeneous ecological degradation. We propose a novel knowledge-guided collaborative bilevel formerpointer reinforcement learning framework (KC-BFPRL) to address this complexity. Using a hierarchical paradigm, KC-BFPRL decomposes RAMP into global task allocation and local restoration planning, with the latter further divided into upper-level trajectory planning and lower-level restoration area allocation. Our specialized architecture pairs featuring a Transformer-based encoder that fuses static environmental features with dynamic UAV states, and a Pointer Network decoder trained via a robust actor-critic framework. By embedding ecological priority rules and heuristic logic, KC-BFPRL achieves a structured warm-start, solving the RL cold-start problem while ensuring strict constraint satisfaction. Extensive experiments demonstrate that KC-BFPRL consistently outperforms state-of-the-art baselines, achieving superior objective values and efficiency. It maintains a $0.00\%$ optimality gap in the most complex scenarios U8-R160 and operates nearly three times faster than MAPDP, validating its robustness, scalability, and real-time applicability for large-scale automated ecological restoration.

↑