arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动力学信息强化学习实现单足跳跃四旋翼的敏捷与节能运动

Dynamics-Informed Reinforcement Learning for Agile and Energy-Efficient Locomotion of a Monopedal Hopping Quadcopter

Ruigang Chen, Qi Zhang, Zhicheng Zhong, Zhuorui Yun, Yizhar Or, Mingyi Liu

arXiv 2609.15399首次发表:更新:

发表机构

Guangdong Technion - Israel Institute of Technology; Technion-Israel Institute of Technology(广东以色列理工学院; 以色列理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出动力学信息强化学习框架,通过嵌入比能和相位奖励,实现单足跳跃四旋翼的稳定高效跳跃,能耗较悬停和低效基线分别降低82%和73%。

AI 中文摘要

尽管空中-腿部机器人兼具敏捷性和效率,但在复杂混合动力学下控制高速跳跃仍具挑战性。强化学习(RL)前景广阔,但容易产生低能效的“奖励黑客”行为。我们提出了一种用于单足跳跃四旋翼的动力学信息强化学习框架。通过将目标比能嵌入奖励函数,我们将优化约束在物理可行的能量流形上,确保稳定的跳跃行为。通过奖励相位一致行为,可鼓励仿生支撑相冲量。此外,惩罚机电功率浪费促使电机产生高效冲量。这使得策略能够在无启发式状态机的情况下,严格在弹簧恢复期间注入能量。MuJoCo模拟验证了在严重姿态-接触耦合下,高度调节和前进速度跟踪(高达2.0米/秒)的鲁棒性。最终,我们的方法产生了高度敏捷的跳跃步态,与悬停基线和低效基线相比,能耗分别降低了82%和73%。

英文摘要

Although aerial-legged robots offer combined agility and efficiency, controlling high-speed hopping under complex hybrid dynamics is challenging. Reinforcement Learning (RL) is promising but prone to energy-inefficient "reward hacking". We propose a Dynamics-Informed RL framework for a monopedal hopping quadcopter. By embedding a target Specific Energy into the reward, we constrain the optimization to a physically viable energy manifold, ensuring stable hopping behaviour. By rewarding the phase-consistent behavior, it can encourage bio-inspired stance-phase impulse. Furthermore, penalizing the electro-mechanical power waste induces the motors generate an efficient impulse. This enables the policy to inject energy strictly during spring restitution without heuristic state machines. MuJoCo simulations validate robust height regulation and forward velocity tracking up to 2.0 m/s despite severe attitude-contact coupling. Ultimately, our approach yields a highly agile hopping gait, reducing energy consumption by 82% and 73% compared to hovering baselines and inefficiency baseline, respectively.

Comments9 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑