arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FERPO:前向熵正则化策略优化

FERPO: Forward Entropy-Regularized Policy Optimization

Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

arXiv 2610.02198首次发表:更新:

发表机构

Applied and Theoretical Aspects of Robot Intelligence (ATARI) Lab; Munich Institute of Robotics and Machine Intelligence (MIRMI); Technical University of Munich(应用与理论机器人智能实验室; 慕尼黑机器人与机器智能研究所; 慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对评论家动作梯度不可靠的问题,提出前向熵正则化策略优化(FERPO),利用前向KL目标与自归一化重要性采样改进策略,促进探索,在MuJoCo和ManiSkill上表现优异且更新更快。

AI 中文摘要

连续控制中的在线强化学习,几种最先进的方法利用学习到的评论家的动作梯度来改进策略。然而,评论家通常被训练来预测回报,而准确的价值预测并不一定能产生准确的动作导数,这可能导致不可靠的策略更新。我们提出了前向熵正则化策略优化(FERPO),一种在策略的最大熵强化学习算法,它利用评论家值进行策略改进,而无需对评论家关于动作进行微分。FERPO 从由熵和 Kullback-Leibler(KL)散度正则化的策略改进目标中推导出最优目标动作分布。然后,我们通过最小化前向 KL 目标来将演员拟合到该目标,该目标使用从 rollout 策略中抽取的动作的自归一化重要性采样(SNIS)进行估计。通过限制目标分布与 rollout 策略的偏差,KL 正则化有助于保持这些重要性权重的良好行为。与可能偏向目标分布某些模态的反向 KL 目标相反,前向 KL 目标鼓励覆盖多个高价值模态,从而促进探索。在 MuJoCo Playground 和 ManiSkill 上的实验和消融研究显示了有竞争力的性能和样本效率提升。计算基准测试也证明了比相对熵路径策略优化(REPPO)更快的演员更新。

英文摘要

Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).

CommentsCode: https://github.com/Atarilab/FERPO

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑