去中心化主脑:多智能体路径规划中通过迭代意图去噪的联合动作精化
Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
浏览论文内容
中文总结 AI 辅助
针对去中心化多智能体路径规划中独立采样导致联合动作不兼容的问题,提出DMM方法,通过跨通信轮次的迭代意图去噪精化动作,结合模仿学习预训练与MICPO强化学习优化,在1,600个MovingAI任务中解决1,598个,并支持百万级智能体扩展。
中文摘要 AI 辅助
带通信的去中心化多智能体路径规划(MAPF)要求智能体在部分可观测条件下无碰撞地到达各自目标。基于专家数据训练的可学习策略为此问题提供了有效方法。然而,当多个协调的联合动作在同一上下文中均有效时,从各智能体分布中独立采样可能会将局部有效的选择重新组合成不兼容的联合动作。即使各智能体动作分布被正确学习,这种失败也可能源于最终的采样机制。DMM(去中心化主脑)通过将一次性动作采样替换为跨通信轮次的离散、迭代式动作意图精化来解决此问题,其灵感来自扩散模型中的去噪过程。智能体初始化随机动作意图,并通过局部通信进行精化,在承诺前耦合各自的选择。DMM使用模仿学习在专家MAPF解决方案上进行预训练,并进一步通过MICPO(一种专为多智能体、多轮动作精化设计的无评论家群体相对强化学习方法)进行优化。DMM通常比所评估的可学习基线实现更高的成功率和更低的解决方案成本。在1,600个MovingAI任务上,使用MICPO微调的DMM解决了1,598个,是所评估方法中覆盖率最高的,同时其解决方案成本接近最强基线的水平。DMM还能在障碍密集环境中扩展到超过一百万个同时行动的智能体。这些结果表明,轮级意图精化可以改善联合动作协调,同时保持去中心化执行。
英文摘要
Decentralized multi-agent path finding (MAPF) with communication requires agents to reach individual goals without collisions under partial observability. Learnable policies trained on expert data provide an effective approach to this problem. However, when several coordinated joint actions are valid in the same context, independently sampling from per-agent distributions can recombine locally valid choices into incompatible joint actions. This failure can arise from the final sampling mechanism even when the per-agent action distributions are learned correctly. DMM (Decentralized Master-Mind) addresses this by replacing one-shot action sampling with discrete, iterative refinement of action intents across communication rounds, inspired by denoising in diffusion models. Agents initialize random action intents and refine them through local communication, coupling their choices before commitment. DMM is pretrained with imitation learning on expert MAPF solutions and further optimized with MICPO, a critic-free group-relative reinforcement-learning method designed for multi-agent, multi-round action refinement. DMM generally achieves higher success rates and lower solution costs than the evaluated learnable baselines. On 1,600 MovingAI tasks, DMM fine-tuned with MICPO solves 1,598, the highest coverage among the evaluated methods, while achieving solution costs close to those of the strongest baselines. DMM also scales to over one million simultaneously acting agents in obstacle-rich environments. These results show that round-level intent refinement can improve joint-action coordination while preserving decentralized execution.