CHORUS: 基于单一VLA策略的去中心化多体协作
CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
浏览论文内容
中文总结 AI 辅助
提出CHORUS框架,利用预训练视觉-语言-动作模型的视觉运动先验,实现无需推理时通信的去中心化多机器人协作,在真实实验中显著优于基线。
中文摘要 AI 辅助
多机器人协作使机器人能够高效完成从通过门搬运沙发到建筑工地组装结构等各种任务。然而,在移动多机器人环境中实现这种协调仍然具有挑战性:基于团队联合观测的集中式方法随团队规模扩展性差,而为每个机器人训练一个策略的去中心化方法通常需要显式对齐程序或推理时信息共享来克服部分可观测性。我们的关键见解是,预训练的视觉-语言-动作(VLA)模型的视觉运动先验应能够仅从每个机器人的局部观测实现反应式去中心化协作,无需这些推理时假设。我们提出CHORUS,一个适配单一VLA骨干以控制多样化多机器人团队的框架。推理时,每个机器人运行CHORUS的独立副本,仅基于其自身观测和机器人标识提示。在包括移动卷尺测量、图书馆书籍交接和洗衣篮抬举的真实实验中,CHORUS相比去中心化从头训练模型提升64个百分点,对队友行为的反应性提升40个百分点,并优于集中式基线。这些结果表明,共享VLA骨干能够实现去中心化多机器人协作,无需每个机器人的独立策略或推理时机器人间通信。
英文摘要
Multi-robot collaboration allows robots to efficiently take on a wide range of tasks, from moving a couch through a doorway to assembling structures on a construction site. However, achieving such coordination in mobile multi-robot settings remains challenging: centralized methods conditioned on the combined observations of a team scale poorly with team size, and decentralized methods that train one policy per robot often require explicit alignment procedures or information sharing at inference time to overcome partial observability. Our key insight is that the visuomotor priors of pretrained vision-language-action (VLA) models should enable reactive, decentralized collaboration from each robot's local observations alone, without these inference-time assumptions. We propose CHORUS, a framework that adapts a single VLA backbone to control diverse, multi-robot teams. At inference time, each robot runs an independent copy of CHORUS, conditioned only on its own observations and a robot-identifying prompt. In real-world experiments including mobile tape measurement, library book handovers, and laundry basket lifting, CHORUS achieves a 64% point improvement over decentralized, from-scratch models, improves reactivity to teammate behavior by 40% points, and outperforms centralized baselines. Together, these results show that a shared VLA backbone is capable of achieving decentralized multi-robot collaboration, without per-robot policies or inter-robot communication at inference.
发表机构
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。