arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04905cs.RO

PRIMAL3:基于强化学习与模仿的多智能体路径规划——利用LaCAM3

PRIMAL3: Pathfinding via Reinforcement and Imitation Multi-Agent Learning - Leveraging LaCAM3

  • National University of Singapore(新加坡国立大学)
  • Stanford University(斯坦福大学)
  • Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Chengyang He, Tanishq Duhan, Gadiel Sznaier Camps, Fangyuan Wang, Yuhong Cao, Jiankai Sun, Ge Sun, Mac Schwager, Guillaume Sartoretti

AI总结:

PRIMAL3是整合强化学习、LaCAM3等的超大规模多智能体路径规划框架,性能优于现有基线,可扩展至10万智能体,能部署于物理机器人系统

AI中文摘要:

我们提出PRIMAL3,一种超大规模的基于学习的多智能体路径规划(MAPF)框架,整合了强化学习、拓扑感知通信、LaCAM3引导训练以及基于PIBT的动作优化。PRIMAL3针对拓扑关键状态下的故障,此时智能体必须围绕瓶颈、死胡同和持续冲突进行果断协调。每个智能体用从割点、死胡同区域、最短路径距离和阻塞估计中提取的特征表示。两个互补图捕捉智能体交互:同向跟随图沿兼容路径传播多跳上下文,反向冲突图通过掩码注意力和相对特征区分争夺共享空间的智能体。训练期间,我们提出用策略熵识别不确定智能体,LaCAM3为这些智能体提供置信度触发的动作干预和标签平滑的模仿目标。执行期间,优先级感知的PIBT模块结合持久、学习到的和距离感知的优先级,以及策略感知的弃权(不执行)偏好优化提出的联合动作,同时保持无碰撞执行。该框架将学习到的探索与结构化专家指导相结合,推理时无需LaCAM3。实验表明,PRIMAL3显著优于最先进的基于学习的基线,可扩展到多达城市级100000个智能体的超大规模实例。实际实验进一步证明了在物理机器人系统上部署PRIMAL3的可行性,消融研究验证了我们提出的各组件的单独贡献。项目页面:this https URL

英文摘要:

We present PRIMAL3, an ultra-large-scale learning-based framework for multi-agent pathfinding (MAPF) that integrates reinforcement learning, topology-aware communication, LaCAM3-guided training, and PIBT-based action refinement. PRIMAL3 targets failures at topologically critical states, where agents must coordinate decisively around bottlenecks, dead ends, and persistent conflicts. Each agent is represented using features derived from cut vertices, dead-end regions, shortest-path distances, and blocking estimates. Two complementary graphs capture agent interactions: a same-direction following graph propagates multihop context along compatible paths, while a different-direction conflict graph differentiates agents competing for shared space through masked attention and relative features. During training, we propose to let policy entropy identify uncertain agents, for which LaCAM3 provides confidence-triggered action interventions and label-smoothed imitation targets. During execution, a priority-aware PIBT module refines the proposed joint actions using persistent, learned, and distance-aware priorities together with policy-aware fallback preferences while maintaining collision-free execution. The resulting framework combines learned exploration with structured expert guidance without requiring LaCAM3 at inference. Experiments demonstrate that PRIMAL3 substantially outperforms state-of-the-art learning-based baselines and scales to ultra-large instances with up to city-level 100,000 agents. Real-world experiments further demonstrate the feasibility of deploying PRIMAL3 on physical robotic systems and ablation studies validate the individual contributions the components we proposed. Project page: https://marmotlab.github.io/PRIMAL3/

补充信息

↑