用于协调多智能体探索的规划器条件扩散模型
Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration
浏览论文内容
中文总结 AI 辅助
提出规划器条件扩散策略PCDP,通过规划器身份作为条件输入训练多模态单智能体策略,结合局部重排序实现协调,在多智能体探索中提升性能并验证了方法有效性。
中文摘要 AI 辅助
协调多智能体探索不仅需要高效的个体覆盖,还要求智能体在较长规划周期内实现非冗余覆盖。传统方法依赖手工设计的协调规则,而端到端多智能体学习方法难以扩展和训练。基于扩散的规划器(如DARE)通过生成长周期轨迹而非单步动作提供了有前景的替代方案,但现有方法在狭窄的规划器分布上训练,限制了行为多样性和推理时的可控性。我们提出了一种用于基于图的多智能体探索的规划器条件扩散策略(PCDP)。PCDP在多种规划器风格的演示上进行训练,将规划器身份作为显式条件输入,使单个共享模型能够学习多模态轨迹分布,并从相同观测中生成多样、可控的轨迹候选。我们未端到端学习协调,而是在所有智能体间复用这种多模态单智能体策略,并通过局部重排序引入协调,即附近智能体共同选择预测重叠最小的轨迹组合。我们在4智能体模拟环境中,针对100个保留地图,将PCDP与经典方法和基于扩散的基线进行评估。PCDP达到了基于扩散的基线的完美成功率,同时改善了平均最大智能体行程、团队总行程和智能体不平衡度。关键的是,仅对单规划器基线进行重排序仅产生微小增益,表明规划器条件多模态是协调改进的主要贡献因素。定性模拟结果和2智能体的真实机器人实验进一步验证,多样的长周期轨迹生成能在无任何显式排斥机制的情况下,使智能体产生自发空间分离。
英文摘要
Coordinated multi-agent exploration requires not only efficient individual coverage but also non-redundant coverage across agents over extended planning horizons. Conventional approaches rely on hand-crafted coordination rules, while end-to-end multi-agent learning methods are difficult to scale and train. Diffusion-based planners such as DARE offer a promising alternative by generating long-horizon trajectories instead of single-step actions, but existing methods are trained on a narrow planner distribution, limiting behavioral diversity and inference-time controllability. We propose a Planner-Conditioned Diffusion Policy (PCDP) for graph-based multi-agent exploration. PCDP is trained on demonstrations from multiple planner styles with planner identity as an explicit conditioning input, enabling a single shared model to learn a multimodal trajectory distribution and generate diverse, controllable trajectory candidates from the same observation. Rather than learning coordination end-to-end, we reuse this multimodal single-agent policy across all agents and introduce coordination through local reranking, in which nearby agents jointly select the trajectory combination with minimal predicted overlap. We evaluate PCDP against classical and diffusion-based baselines on 100 held-out maps in a four-agent simulation setting. PCDP matches the perfect success rate of the diffusion-based baselines while improving mean max-agent travel, total team travel, and agent imbalance. Crucially, reranking alone over a single-planner baseline yields only marginal gains, indicating that planner-conditioned multimodality is the main contributor to improved coordination. Qualitative simulation results and real-robot experiments with two agents further validate that diverse long-horizon trajectory generation produces emergent spatial separation between agents without any explicit repulsion mechanism.
发表机构
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。