arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

残差去噪实现按需的样本高效多智能体协调

Residual Denoising Enables Sample-Efficient Multi-Agent Coordination on Demand

Dayi Dong, Maulik Bhatt, Aayushi Shrivastava, Lasse Peters, Negar Mehr

arXiv 2609.32129首次发表:更新:

AI 中文总结

本文提出ALTER方法,通过残差去噪器将预训练单智能体扩散策略适配为多智能体协调,实现按需协调,同时保持单智能体技能,实验证明其协调成功率和技能保持均优于基线。

AI 中文摘要

预训练的机器人策略提供了强大的操作技能,但通常仅限于单智能体设置,即机器人独立行动。在本工作中,我们研究了如何利用最少的协作数据将预训练的单智能体扩散策略适配到多智能体设置,同时优化两个关键目标:高协调性能和单智能体技能保持。为此,我们引入了ALTER,一种按需协调的适配方法:适配后的策略在团队部署时与其他机器人协调,而在单独操作时仍能独立行动。执行是分布式的:每个机器人仅依据自身的视觉观察行动,无需智能体间的显式通信。我们的方法训练一个协调头,该协调头预测一个残差去噪器,在必要时将单智能体行为转换为协调的多智能体行为,同时保持单智能体能力。为了保持单智能体能力,我们在训练残差去噪器时,用基础策略生成的自我蒸馏数据扩充少量协作演示。在仿真中,ALTER在协调成功率上优于我们的基线,同时保持了更高的源技能保持率。在我们的硬件实验中,我们发现了类似的趋势,即ALTER在协调成功率和单智能体技能保持的联合优化上优于基线。

英文摘要

Pretrained robot policies offer strong manipulation skills but are typically limited to single-agent settings, where a robot acts in isolation. In this work, we study how to adapt pretrained single-agent diffusion policies to multi-agent settings using minimal collaborative data, co-optimizing for two key objectives: high coordination performance and single-agent skill retention. To this end, we introduce ALTER, an adaptation method for coordination on demand: the adapted policy coordinates with other robots when deployed in a team while remaining capable of acting independently when operating alone. Execution is decentralized: each robot acts only on its own visual observations, without explicit inter-agent communication. Our method trains a coordination head that predicts a residual denoiser to transform single-agent behavior into coordinated multi-agent behavior when necessary while also preserving single-agent capabilities. To preserve single-agent capabilities, we augment a small number of collaborative demonstrations with self-distilled data generated by the base policy during training of the residual denoiser. In simulation, ALTER achieves higher coordination success over our baselines while retaining much higher source-skill retention. In our hardware experiments, we find similar trends where ALTER better co-optimizes for coordination success and single-agent skill retention than the baselines.

Comments8 pages, 4 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑