arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于强化学习的自行车杂技

Bicycle Acrobatics with Reinforcement Learning

Shamel Fahmi, Arianna Ilvonen, Samuel Zapolsky, Yu-Ming Chen, Ravi Boggavarapu, Paul Drews, Ashwin Khadke, Dean Molinaro, Kaiyu Zheng, Alfred Rizzi, Gabriel Nelson

arXiv 2608.00880首次发表:更新:

AI 中文总结

本文采用强化学习技术,通过引入编排器协调多种RL范式训练的策略,让定制自行车机器人UMV完成多种杂技动作,验证了RL可赋予轮式机器人高敏捷性。

AI 中文摘要

自行车机器人速度快且能效高,但其简单的机械设计、欠驱动与非完整动力学特性,使其难以实现高敏捷性动作。本文采用强化学习(RL)技术,让自行车机器人学习并组合多样化的动态杂技动作库。通过航点跟随、姿态到达、扭转载荷跟踪、引导跟踪及运动模仿等不同RL范式,机器人自主掌握了单跳与多台的前向及侧向跳跃、可转向跳跃、前空翻、起立、下蹲、行驶、翘前轮、兔跳和三点转向等动作。为协调这些行为,我们引入了一个编排器,通过依赖状态的触发器在策略间切换,实现了鲁棒的长程杂技动作。我们在定制自行车机器人Ultra Mobility Vehicle(UMV)的仿真环境与实物平台上验证了该方法:机器人可反复跨越最高1米的台架,在跟随航点时完成15次以上连续自主跳跃,能应对未见过的多台架配置,在20余次连续试验中执行起立、跳跃、空翻、下蹲的连续动作库,还可完成10次以上自主且可转向的翘前轮、侧向跳跃及单轮跳下的动作库。这些结果表明,RL可赋予自行车机器人此前主要与腿式平台相关的敏捷性,同时保留轮式运动的速度与能效,为自行车杂技奠定了基础。

英文摘要

Bicycle robots are fast and energy efficient, but their simple mechanical design and their underactuated and non-holonomic dynamics make highly agile maneuvers difficult to achieve. Here, we use Reinforcement Learning (RL) to enable a bicycle robot to learn and compose a diverse repertoire of dynamic acrobatic stunts. Using different RL formulations such as waypoint following, pose reaching, twist tracking, guided tracking, and motion imitation, the robot acquires autonomous single and multi-table forward and lateral jumps, steerable jumps, front flips, kip-ups, kip-downs, driving, wheelies, bunny hops, and three-point turns. To coordinate these behaviors, we introduce an orchestrator that transitions between policies using state-dependent triggers, enabling robust long-horizon acrobatic stunts. We validate the approach on the Ultra Mobility Vehicle (UMV), a custom bicycle robot, in simulation and hardware. The robot repeatedly traverses tables up to 1 m high, performs more than 15 consecutive autonomous jumps while following waypoints, handles previously unseen multi-table configurations, executes continuous repertoires of kipups, jumps, flips, kip-downs, over more than 20 consecutive trials, and performs more than 10 consecutive autonomous and steerable repertoires of wheelies, lateral jumps, and single-wheel jump downs. These results demonstrate that RL can endow bicycle robots with levels of agility previously associated primarily with legged platforms while preserving the speed and efficiency of wheeled locomotion, establishing a foundation for bicycle acrobatics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑