发表机构
INRIA; CNRS; LORIA; International University of Rabat; ISM(法国国家信息与自动化研究所; 法国国家科学研究中心; 洛林计算机科学与应用实验室; 拉巴特国际大学; 马赛科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对绳驱并联机器人泛化控制难题,提出执行器级策略的深度强化学习方法,可适配任意构型,在鲁棒性、精度上优于传统方法,成功实现跨构型迁移控制。
AI 中文摘要
绳驱并联机器人(CDPR)具有多样的构型与复杂的控制挑战,可通过深度强化学习(DRL)学习其非线性动力学来应对这些挑战。然而,DRL方法通常需要大量训练时间,且生成的策略无法很好地泛化到不同机器人构型或不同数量的执行器。在本文中,我们提出了一种不依赖特定机器人构型的新型CDPR控制DRL方法。与传统DRL方法学习控制整个机器人以达到期望的末端执行器位置不同,我们的方法训练一种执行器级策略(ALP),用于控制每个电机以达到其目标绳索长度。据我们所知,这是首个将DRL应用于采用执行器级策略的CDPR控制的工作。该方法有两个主要优势:(i)单个共享策略可应用于任何CDPR构型,无论执行器数量多少;(ii)无需依赖逆运动学,避免了更具挑战性的正运动学问题。训练在仿真中进行,学习到的策略已成功迁移到真实CDPR。实验结果表明,执行器级策略在鲁棒性和精度上均优于传统强化学习方法。我们进一步用在仿真的4电机平面CDPR(2D运行)上训练的策略,控制了真实的8电机CDPR实现3D运动,这说明所提方法适用于任何CDPR构型,与执行器数量或布置无关。
英文摘要
Cable-driven parallel robots (CDPRs) present diverse configurations and complex control challenges, which can be addressed by deep reinforcement learning (DRL) by learning their nonlinear dynamics. However, DRL methods often require extensive training time, and the resulting policies do not generalize well to different robot configurations or varying numbers of actuators. In this article, we introduce a novel DRL approach for controlling CDPRs that does not depend on the specific robot configuration. Our method trains an actuator-level policy that controls each motor to achieve its target cable length, in contrast to conventional DRL approaches that learn to control the entire robot to reach a desired end-effector position. To the best of our knowledge, this is the first work to apply DRL to control CDPRs using an actuator-level policy. This approach offers two main advantages: (i) a single shared policy can be applied to any CDPR configuration, regardless of actuator count, and (ii) reliance on inverse kinematics, avoiding the more challenging forward kinematics problem. Training is performed in simulation, and the learned policy is successfully transferred to a real CDPR. Experimental results show that the actuator-level policy (ALP) surpasses traditional reinforcement learning methods in both robustness and precision. We further control a real 8-motor CDPR with 3D motion using a policy trained on a simulated 4-motor planar CDPR operating in 2D. This illustrates that the proposed method is applicable to any CDPR configuration, independent of actuator number or placement.