arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于带旋转的鲁棒终身多智能体路径规划的搜索辅助联合智能体-环境强化学习

Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations

He Jiang, Jingtian Yan, Yulun Zhang, Yimin Tang, Tanishq Duhan, Rishi Veerapaneni, Guillaume Sartoretti, Jiaoyang Li

arXiv 2608.05588首次发表:更新:

AI 中文总结

针对带旋转约束的鲁棒终身多智能体路径规划难题,提出搜索辅助联合强化学习方法,通过结合Causal PIBT与联合优化智能体-环境策略,在高密度地图及混合现实仓库环境中取得显著性能提升。

AI 中文摘要

终身多智能体路径规划(LMAPF)需要为智能体反复规划无碰撞路径,这些智能体在完成当前目标后会持续接收新目标。尽管已提出许多基于学习的LMAPF规划器,但多数依赖过于简化的运动学假设,可能忽略对实际性能至关重要的运动约束。本研究针对源自众多现实世界自动化仓库系统的更现实LMAPF模型(称为LMAPF-R2)展开研究,该模型包含鲁棒安全约束和原地旋转约束,这些约束大幅提升了协调难度,尤其在高度受限空间中。为应对这些挑战,我们提出搜索辅助联合强化学习(SJRL):首先为神经策略补充基于单步搜索的规划器Causal PIBT,该规划器可解决智能体碰撞并传播其意图;随后引入统一的RL公式,联合优化智能体策略与环境策略,其中环境策略学习图边代价,通过反向Dijkstra搜索提供全局运动指导。实验表明,在多个高密度地图上,SJRL相比强大的基于搜索的规划器Causal-PIBT实现了显著性能提升;我们还在包含8台物理机器人和248台虚拟机器人的具挑战性的混合现实仓库环境中验证了SJRL的有效性。

英文摘要

Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals upon reaching their current ones. While many learning-based planners have been proposed for LMAPF, most rely on oversimplified kinematic assumptions that may overlook motion constraints critical to real-world performance. In this work, we study a more realistic LMAPF model derived from many real-world automated warehouse systems, termed LMAPF-R2, which incorporates robust safety constraints and in-place rotation constraints. These constraints substantially increase coordination difficulty, particularly in highly constrained spaces. To address these challenges, we propose Search-Aided Joint Reinforcement Learning (SJRL). We first augment neural policies with Causal PIBT, a single-step search-based planner that resolves agents' collisions and propagates their intentions. We then introduce a unified RL formulation that jointly optimizes agent and environment policies, where the environment policy learns graph edge costs to provide global movement guidance via backward Dijkstra search. Experiments demonstrate that SJRL achieves significant improvements over the strong search-based planner, Causal-PIBT, across multiple high-density maps. We further validate SJRL in a challenging mixed-reality warehouse environment with 8 physical robots and 248 virtual robots.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑