arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习工厂中的多智能体任务分配与导航:从仿真到真实机器人

Learning Multi-Agent Task Assignment and Navigation in the Factory: from Simulation to Real Robots

Abdalwhab Bakheet Mohamed Abdalwhab, Giovanni Beltrame, David St-Onge

arXiv 2609.14567首次发表:更新:

发表机构

INIT Robots Lab, École de technologie supérieure; MISTLab, Department of Computer and Software Engineering, Polytechnique Montréal(INIT机器人实验室,高等工程技术学院; MIST实验室,蒙特利尔理工学院计算机与软件工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出特征融合多智能体近端策略优化(FMAPPO),实现工厂中多机器人任务分配与导航的仿真到现实部署,显著提升效率与安全性。

AI 中文摘要

强化学习(RL)在机器人决策方面显示出相当大的潜力,然而在工业环境中将多智能体强化学习(MARL)部署到物理多机器人系统上仍然具有挑战性。本文研究了去中心化MARL在多机器人多机器照料中的实际应用性。我们提出了特征融合多智能体近端策略优化(FMAPPO),该方法融合了2D激光雷达测量与任务特定的状态信息,以实现安全的去中心化多机器人任务分配和导航。我们开发了一个完整的仿真到现实流水线,使用高保真机器人仿真和ROS2,并部署在现实世界条件下运行的物理移动机械臂平台上,实验期间机械臂被禁用。我们进一步研究了所学策略对命令更新频率的敏感性,这是现实世界部署中的一个重要考虑因素。仿真中的对比评估表明,FMAPPO以较大的效应量显著优于最先进的基线,在零件交付方面分别比MAPPO和SMAPPO提高了106%和21%,在零件收集方面分别提高了48%和11%。FMAPPO还分别将机器利用率提高了31和10个百分点,同时与MAPPO和SMAPPO相比,碰撞减少了18%和15%,安全评分分别提高了14和6个百分点。此外,现实世界实验表明,所学的去中心化策略可以在现实世界的感知和控制约束下协调多个机器人为多台机器服务,同时保持安全运行。现实世界实验的视频可在网上获取,网址为https://this URL。

英文摘要

Reinforcement learning (RL) has shown considerable promise for robotic decision-making, yet deploying multi-agent RL (MARL) on physical multi-robot systems in industrial environments remains challenging. This paper investigates the real-world applicability of decentralized MARL for multi-robot multi-machine tending. We propose Feature-fusion Multi-Agent Proximal Policy Optimization (FMAPPO), which fuses 2D LiDAR measurements with task-specific state information to enable safe decentralized multi-robot task assignment and navigation. A complete simulation-to-reality pipeline was developed using high-fidelity robotic simulation and ROS2 and deployed on physical mobile-manipulator platforms operating under realistic real-world conditions, with the robotic arms disabled during the experiments. We further investigate the sensitivity of the learned policy to command update frequency, an important consideration for real-world deployment. Comparative evaluation in simulation demonstrated that FMAPPO significantly outperformed state-of-the-art baselines with a large effect size, achieving improvements of 106\% and 21\% in parts delivery and 48\% and 11\% in parts collection over MAPPO and SMAPPO, respectively. FMAPPO also increased machine utilization by 31 and 10 percentage points, respectively, while reducing collisions by 18\% and 15\% and increasing the safety score by 14 and 6 percentage points compared with MAPPO and SMAPPO, respectively. Furthermore, real-world experiments demonstrated that the learned decentralized policies can coordinate multiple robots to service multiple machines while maintaining safe operation under real-world sensing and control constraints. Videos of the real-world experiment are available online https://anonymouspapers123.github.io/FMAPPO/.

CommentsThis work has been submitted for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑