发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对真人大小双臂机器人的稳健摆臂运动难题,提出路径点引导强化学习(WGRL)方法,结合路径点引导与任务成功、机械能奖励,经仿真与硬件实验验证,实现了含故障恢复的稳健摆臂,为机器人手臂运动设计提供指南。
AI 中文摘要
摆臂运动是灵长类动物主要依靠手臂进行移动的一种运动形式,使其能够在无立足点的环境中移动。然而,这种运动需要高度协调的全身运动以及对横杆抓取和释放的精准时序控制,因此在真人大小的机器人平台上实现稳健的摆臂行为仍然具有挑战性。本研究提出一种基于强化学习的方法,以在真人大小的双臂机器人上实现摆臂运动。所提方法的核心是路径点引导强化学习(Waypoint-Guided Reinforcement Learning, WGRL),这是一种用于诱导非线性复杂运动的学习框架。对于无法获取模仿学习数据的高难度任务,WGRL通过为末端执行器轨迹稀疏指定路径点来引导行为获取,同时通过强化学习生成全身运动。此外,通过将路径点跟随引导与基于任务成功和机械能的奖励相结合,并在为仿真到现实(Sim-to-Real)迁移设计的环境中进行训练,所提方法实现了向前移动和运动稳定性。所获得的行为在具有几何变化的单杠环境中通过仿真到仿真(Sim-to-Sim)实验以及硬件实验进行评估,证实了包括故障恢复行为在内的稳健摆臂运动。本研究为在真人大小的机器人硬件上实现基于手臂的运动以及扩大机器人的可移动工作空间提供了有效的学习设计指南。
英文摘要
Brachiation is a form of locomotion in which primates move primarily using their arms, enabling traversal in environments without footholds. However, this motion requires highly coordinated whole-body movement and precise timing control for bar grasping and release. As a result, achieving robust behavior on life-sized robotic platforms remains challenging. In this study, we present a reinforcement learning-based method to realize brachiation on a life-sized dual-arm robot. The core of the proposed approach is Waypoint-Guided Reinforcement Learning (WGRL), a learning framework for inducing non-linear and complex motions. For high-difficulty tasks where imitation learning data are unavailable, WGRL guides behavior acquisition by sparsely specifying waypoints for the end-effector trajectory, while whole-body motion is generated through reinforcement learning. In addition, by integrating the waypoint-following guidance with rewards based on task success and mechanical energy, and training in an environment designed for Sim-to-Real transfer, the proposed method achieves both forward progression and motion stability. The acquired behavior is evaluated through Sim-to-Sim experiments under monkey-bar environments with geometric variations and hardware experiments, confirming robust brachiation including failure recovery behavior. This study provides effective learning design guidelines for realizing arm-based locomotion on life-sized robotic hardware and expanding the traversable workspace of robots.
CommentsAccepted to 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)