arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09587cs.RO

CoDiMAD:基于扩散的特权蒸馏用于无通信多机器人协调

CoDiMAD: Diffusion-Based Privileged Distillation for Communication-Free Multi-Robot Coordination

发表机构北京大学 · 国防科技创新研究院
查看机构详情
  • Peking University(北京大学)
  • National Innovation Institute of Defense Technology(国防科技创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

Jiyue Tao, Shunheng Xin, Tongsheng Shen, Dexin Zhao, Feitian Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究无通信多机器人协调问题,提出CoDiMAD三阶段框架,先训练特权预言机,再构建数据集,最后将预言机蒸馏到学生智能体中,通过扩散反向过程采样动作,实验证明其性能优于基线。

中文摘要 AI 辅助

在部分可观测性下的分散式多机器人协调仍然具有挑战性,特别是在无通信设置中,智能体必须仅根据局部传感器观测行动。特权策略蒸馏通过将知识从全局信息的预言机转移到受传感器约束的学生智能体提供了一种有前景的方法。但在多智能体系统中,相同局部观测可能对应多种全局配置,导致条件动作分布固有地多模态。标准确定性蒸馏将这些模式合并为均值,常产生无效或犹豫的动作。为解决此问题,我们提出CoDiMAD,一个三阶段框架,用MAPPO训练特权预言机,构建局部观测 - 预言机 - 动作对的离线数据集,并将预言机蒸馏到参数化为条件去噪扩散概率模型的分散式学生智能体中。通过扩散反向过程近似条件预言机 - 动作分布,CoDiMAD从连贯协调模式中采样决定性动作而非求平均。理论分析刻画了确定性蒸馏的模式平均失败和基于扩散蒸馏的分布恢复特性。在三个合作任务上的实验表明CoDiMAD持续优于直接局部MARL和确定性蒸馏基线。源代码将在接受后公开。

英文摘要

Decentralized multi-robot coordination under partial observability remains challenging, especially in communication-free settings where agents must act solely from local sensor observations. Privileged policy distillation provides a promising approach by transferring knowledge from a globally informed oracle to sensor-constrained students. However, in multi-agent systems, the same local observation may correspond to multiple global configurations requiring qualitatively different cooperative actions, making the conditional action distribution inherently multi-modal. Standard deterministic distillation collapses these modes to their mean, often yielding invalid or hesitant actions. To address this issue, we propose CoDiMAD, a three-stage framework that trains a privileged oracle with MAPPO, constructs an offline dataset of local-observation-oracle-action pairs, and distills the oracle into decentralized students parameterized as conditional denoising diffusion probabilistic models. By approximating the conditional oracle-action distribution through the diffusion reverse process, CoDiMAD samples decisive actions from coherent coordination modes rather than averaging across them. Theoretical analysis characterizes the mode-averaging failure of deterministic distillation and the distributional recovery property of diffusion-based distillation. Experiments on three cooperative tasks show that CoDiMAD consistently outperforms direct local MARL and deterministic distillation baselines. The source code will be made publicly available upon acceptance.

↑