arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DuoMind:通过语义通信实现分布式多机器人协调

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang

arXiv 2610.02161首次发表:更新:

发表机构

University of California, Davis; Microsoft Research; Analog Devices(加州大学戴维斯分校; 微软研究院; 亚德诺半导体)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DuoMind分布式分层框架,通过VLM编排器与VLA动作模型结合语义通信实现多机器人长期任务协调,并构建RoboPoly基准验证性能提升。

AI 中文摘要

视觉语言模型(VLM)和视觉语言动作模型(VLA)近期推动了通用机器人的快速发展,但大多数进展集中在单机器人场景。将这些能力扩展到多机器人系统仍然具有挑战性,因为机器人必须在保持可靠、细粒度执行的同时协调长期行为。我们提出了DuoMind,一个通过语义通信实现多机器人协调的分布式分层框架。每个机器人使用基于VLA的动作模型进行低级执行,并使用基于VLM的编排器进行高级推理和智能体间协调。在每个规划步骤中,每个机器人的编排器对任务指令、局部观测以及从其他机器人接收的消息进行推理,然后为动作模型生成低级指令,并为对等机器人生成语义消息。该架构通过结合VLM的语义推理能力与VLA的精确动作生成能力,利用了预训练模型的互补优势。为解决多机器人协调基准稀缺的问题,我们进一步开发了RoboPoly,一个包含需要分布式控制下协调、闭环执行的长期操作任务的基准。在RoboPoly和RoboTwin上的实验表明,DuoMind提高了多机器人任务性能,而消融研究确认了分层编排和语义通信的贡献。更多细节可在我们的项目页面上获取。

英文摘要

Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑