学习通信条件生成策略用于分散式多智能体碰撞避免
Learning Communication-Conditioned Generative Policies for Decentralized Multi-Agent Collision Avoidance
浏览论文内容
中文总结 AI 辅助
提出分散式通信条件生成框架,利用流匹配策略从离线演示学习,通过潜在消息协调多智能体碰撞避免,实现接近专家性能并泛化至密集场景和真实机器人。
中文摘要 AI 辅助
在这项工作中,我们提出了一种用于多智能体碰撞避免的分散式通信条件生成框架。智能体使用流匹配策略生成短时域动作序列,该策略从可访问全局状态的特权离线演示中训练得到。演示不包含显式通信信号;相反,智能体学习在部分可观测性下交换和聚合编码交互相关意图的潜在消息。这种表述支持测试时的灵活推理,其中无条件生成对应独立行为,而通信条件生成使得无需集中规划即可实现协调交互。所得策略在执行时以完全分散的方式运行,仅依赖局部观测和学习到的消息。结合滚动时域推理方案,所提方法能够高效地单步推理短时域动作序列,并在通信中断时优雅降级。大量仿真结果表明,该方法实现了接近专家的碰撞避免性能,对更密集、未见过的多智能体场景具有强泛化能力,并实现了向真实机器人实验的零样本迁移。
英文摘要
In this work, we propose a decentralized communication-conditioned generative framework for multi-agent collision avoidance. Agents generate short-horizon action sequences using a flow-matching policy trained from privileged offline demonstrations with access to global state. The demonstrations do not include explicit communication signals; instead, agents learn to exchange and aggregate latent messages that encode interaction-relevant intent under partial observability. This formulation supports flexible inference at test time, where unconditioned generation corresponds to independent behavior and communication-conditioned generation enables coordinated interaction without centralized planning. The resulting policies operate in a fully decentralized manner at execution time, relying only on local observations and learned messages. Combined with a receding-horizon inference scheme, the proposed approach enables efficient single-step inference of short-horizon action sequences and degrades gracefully under communication dropouts. Extensive simulation results demonstrate near-expert collision avoidance performance and strong generalization to denser, unseen multi-agent scenarios, along with zero-shot transfer to real-robot experiments.
发表机构
- Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。