AI 中文总结
本文提出D-VLC框架,结合去中心化异步推理等技术,让通用VLM生成机器人特定动作,多场景实验显示其任务成功率超70%,完成时间较几何贪心基线最多缩短55.8%。
AI 中文摘要
多机器人系统,尤其是异构机器人群,可通过并行协作与互补能力提升复杂任务执行效率。但传统基于规则的方法依赖预定义任务模型与专用决策程序,难以理解复杂语义指令并协调异构机器人。大语言模型(LLM)具备强大的语言理解与任务推理能力,可使多机器人系统解读指令、分解任务并依据任务语义分配角色;视觉-语言模型(VLM)进一步整合视觉感知,让机器人能推理物理环境中的对象、区域与空间关系。然而,现有基于LLM/VLM的方法常依赖已知地图、集中式同步决策,限制了其对异构机器人与未见过任务的泛化性。为此,本文提出一种结合去中心化异步推理、轻量级信息共享、能力感知协作及统一动作接口的框架,使通用VLM生成机器人特定动作,由免学习的专家执行,无需针对任务或机器人进行专门训练。在多种场景与多个VLM上开展的实验显示,该方法的成功率超过70%,与几何贪心基线相比,完成时间最多缩短55.8%。
英文摘要
Multi-robot systems, particularly heterogeneous robot swarms, can improve the efficiency of complex task execution through parallel collaboration and complementary capabilities. However, conventional rule-based methods rely on predefined task models and specialized decision making programs, making it difficult to understand complex semantic instructions and coordinate heterogeneous robots. LLMs introduce strong language understanding and task reasoning capabilities, allowing multi-robot systems to interpret instructions, decompose tasks, and assign roles according to task semantics. VLMs further incorporate visual perception, enabling robots to reason about objects, regions, and spatial relationships in physical environments. Nevertheless, existing LLM/VLM based methods often depend on known maps, centralized and synchronized decision making, limiting their generalization to heterogeneous robots and unseen tasks. We therefore propose a framework that combines decentralized asynchronous reasoning, lightweight information sharing, capability aware collaboration, and a unified action interface, enabling general purpose VLMs to generate robot specific actions executed by learning free experts without task or robot specific training. Experiments across diverse scenarios and multiple VLMs show success rates above 70\%, with completion time reduced by up to 55.8\% relative to the geometric greedy baseline.