李雅普诺夫指数作为物理信息密集奖励:强化学习发现超越卡皮察摆的稳定方法
Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza Pendulum
AI总结:
该研究针对强化学习中稳定垂直运动倒立摆的问题,提出将李雅普诺夫特征指数作为密集奖励信号,智能体借此成功找到卡皮察摆振荡运动并抑制摆动,让摆处于直立位置。
AI中文摘要:
我们建议将李雅普诺夫特征指数(LCE)用作强化学习问题中稳定垂直运动倒立摆的密集奖励信号。借助LCE,智能体不仅成功找到了卡皮察摆的振荡运动,还抑制了摆的摆动,使其处于严格直立位置。
英文摘要:
We suggest using the Lyapunov characteristic exponent (LCE) as a dense reward signal for the reinforcement learning problem of stabilizing the inverted pendulum with vertical motion. With LCE, the agent not only successfully found the oscillatory motion known as the Kapitza pendulum but also damped the pendulum's pivoting, leaving it in a strictly upright position.