arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAST:交替状态值目标与扩展策略梯度用于基于模型的强化学习

CAST: Alternating State-Value Targets and Expanded Policy Gradients for Model-Based Reinforcement Learning

Pietro Noah Crestaz, Mohamed Yassine Kabouri, Nicolas Mansard, Andrea Del Prete

arXiv 2609.08853首次发表:更新:

发表机构

University of Trento; LAAS-CNRS; Université de Toulouse; CNRS; New York University; Artificial and Natural Intelligence Toulouse Institute (ANITI)(特伦托大学; 法国国家科学研究中心LAAS实验室; 图卢兹大学; 法国国家科学研究中心; 纽约大学; 图卢兹人工与自然智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CAST提出交替状态值目标与扩展策略梯度,利用规划器引导行为改进价值学习,在多个基准和真实四足机器人上验证了有效性。

AI 中文摘要

基于模型的强化学习(MBRL)是一类学习环境模型并利用该模型进行动作选择的强化学习方法,由于其样本效率高,非常适合机器人技术。将学习到的模型与在线规划相结合可以进一步改善动作选择,因为规划器可以利用模型找到比单独的学习策略更好的动作。最近将学习策略与在线规划相结合的方法通常学习策略的价值,而不是更强的规划器引导行为。我们提出了CAST(交替状态值目标的评论家),它利用规划器引导行为来改进价值学习,同时用当前策略对价值估计进行正则化。CAST用状态值评论家取代了动作值评论家,其训练目标结合了真实的规划器引导转移和当前策略下的想象转移。由此产生的价值函数对应于规划器引导行为与当前策略之间的交替过程,使其能够从更强的规划器行为中受益,同时受到正在学习的策略的正则化。我们在DeepMind Control和HumanoidBench套件上将CAST与几种最先进的方法进行了评估,并展示了在物理Unitree Go2四足机器人上执行动态倒立动作的成功迁移。

英文摘要

Model-based reinforcement learning (MBRL) is a family of RL methods that learn a model of the environment and use it for action selection, making it well suited to robotics due to its sample efficiency. Combining learned models with online planning can further improve action selection, as the planner can exploit the model to find better actions than the learned policy alone. Recent methods combining learned policies with online planning typically learn the value of the policy rather than the stronger planner-guided behavior. We present CAST (Critic with Alternating State-value Target), which uses planner-guided behavior to improve value learning while regularizing the value estimate with the current policy. CAST replaces the action-value critic with a state-value critic, trained using a target that combines a real planner-guided transition and an imagined transition under the current policy. The resulting value function corresponds to an alternating process between planner-guided behavior and the current policy, allowing it to benefit from the stronger planner behavior while being regularised by the policy being learned. We evaluate CAST on the DeepMind Control and HumanoidBench Suites against several state-of-the-art methods, and demonstrate successful transfer to a physical Unitree Go2 quadruped performing a dynamic handstand.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑