发表机构
University of Cambridge; Technical University of Berlin; National Institute of Advanced Industrial Science and Technology (AIST)(剑桥大学; 柏林工业大学; 产业技术综合研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GAMBIT 通过模仿学习和强化学习学习协调的运动原语选择,并引入带备份轨迹的安全机制,实现连续域中上千个机器人的无碰撞轨迹规划,显著优于现有基线。
AI 中文摘要
GAMBIT是国际象棋中的一种开局走法,棋手牺牲一个棋子(通常是兵)以在后续对局中获得位置优势。类似地,在多机器人协调中,单个机器人可能需要放弃局部奖励最大化的行为,以提升整体团队性能。这种自我牺牲的行为难以通过手动设计的启发式方法捕获,尤其是在密集、交互丰富的环境中。本研究聚焦于双积分器连续动力学,探讨如何学习基于运动原语的协调启发式方法,用于多机器人轨迹执行。我们的框架GAMBIT首先通过模仿学习学习协调的运动原语选择,随后通过强化学习对策略进行微调。我们进一步引入了一种带备份轨迹的安全 rollout 机制,保证执行过程中始终无碰撞。实验表明,GAMBIT 显著优于一系列基线方法,包括集中式运动规划器和分散式反应式规划器,并展现出强大的可扩展性。特别是,在连续域中,它能够协调超过一千个机器人,规划延迟低于几百毫秒。
英文摘要
GAMBIT is an opening chess move in which a player sacrifices a piece, typically a pawn, to gain a positional advantage later in the game. Analogously, in multi-robot coordination, individual robots may need to forgo locally reward-maximising behaviours to improve overall team performance. Such self-sacrificial behaviours are difficult to capture with manually designed heuristics, particularly in dense, interaction-rich environments. Focusing on double-integrator continuous dynamics, this work studies how to learn such coordinated heuristics over motion primitives for multi-robot trajectory execution. Our framework, GAMBIT, first learns coordinated motion-primitive selection through imitation learning and subsequently fine-tunes the policy through reinforcement learning. We further introduce a safeguarded rollout mechanism with backup trajectories that guarantees collision-free execution at all times. Experiments demonstrate that GAMBIT substantially outperforms a range of baselines, including centralised motion planners and decentralised reactive planners, while exhibiting strong scalability. In particular, it coordinates over a thousand robots with planning latency below a few hundred milliseconds in continuous domains.