arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Bellman 遇见 Lyapunov:通过掌控混沌实现无监督强化学习

Bellman Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos

Tristan Shah, Wooyoung Chung, Volodomyr Makarenko, Juan Wachs, Stas Tiomkin

arXiv 2610.02012首次发表:更新:

发表机构

Texas Tech University; Purdue University(德克萨斯理工大学; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 Forward CIP(F-CIP),一种仅由系统动力学定义、无需选择信息变量的无监督强化学习目标,可无监督发现平衡等原始行为,并结合简单奖励产生跳跃、奔跑等协调步态。

AI 中文摘要

强化学习(RL)是训练智能体的强大范式,然而其成功依赖于人类工程师为每个新任务设计信息丰富的奖励信号的专业知识。无监督强化学习旨在通过内在动机(IM)来减少这种工程负担:奖励信号源自智能体与环境交互本身。然而,现有的内在动机目标涉及信息变量的选择,这重新引入了该领域试图消除的领域专业知识。我们引入了 Forward CIP(F-CIP),这是可控信息生产(CIP)目标的一种基于强化学习的原生表述,该目标仅由系统动力学定义,无需此类选择。我们证明了 F-CIP 与强化学习兼容,并展示了其与现有算法的有效性。使用 F-CIP 训练智能体能够无监督地发现诸如平衡和维持可控性等原始行为,这些行为对于更复杂的机器人行为至关重要。与简单的前向速度奖励相结合,我们的方法能够产生协调的步态,如跳跃和奔跑,而这些行为在其他情况下需要奖励工程才能学习。

英文摘要

Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engineering with intrinsic motivation (IM): reward signals that emerge from the agent environment interaction itself. Existing IM objectives, however, involve the selection of information variables, which re-introduces domain expertise the field has sought to eliminate. We introduce Forward CIP (F-CIP), an RL-native formulation of the Controllable Information Production (CIP) objective, which is defined by the system's dynamics alone and requires no such selection. We prove that F-CIP is compatible with RL and demonstrate its effectiveness with existing algorithms. Training agents with F-CIP results in unsupervised discovery of primitive behaviors such as balancing and maintaining controllability, which are essential for more complex robot behaviors. Paired with a simple forward-velocity reward, our method produces coordinated gaits such as hopping and running which otherwise require reward engineering to learn.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑