REFINEPPO:通过迭代动作细化学习连续控制策略
REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出REFINEPPO,通过迭代动作细化机制让策略逐步改进动作,结合PPO算法,在14个基准任务上达到或超越标准PPO,并加快收敛。
AI中文摘要:
深度强化学习(DRL)已在广泛的连续控制问题中取得了强劲的性能。然而,这些连续控制策略通常被定义为从观测状态到动作或动作分布的直接映射,需要单个前馈网络在一次前向传播中构建最优控制决策。虽然这种方法有效,但一旦形成初始预测,策略几乎没有机会重新考虑或逐步改进动作。在这项工作中,我们探索了一种替代方法:策略能否不仅学习直接预测动作,还能学习迭代地改进动作,并且这种迭代过程能否在策略学习期间提供优势?我们引入了迭代动作细化(IAR),一种迭代动作构建方法,通过一系列学习到的残差修正来构建控制动作。从初始提议开始,一个共享的细化网络反复以观测状态和当前动作提议为条件,使每个细化步骤都能修正先前步骤构建的动作。最终细化后的提议被用于确定智能体执行的动作。我们将这种迭代动作构建机制与近端策略优化(PPO)相结合,产生了REFINEPPO。我们在14个基准控制任务上评估了REFINEPPO,并辅以对细化深度和更新计划的受控消融实验,以及旨在理解迭代细化为何有效的分析。在这些环境中,REFINEPPO匹配或超过了标准PPO的性能,同时在多个任务上表现出更快的收敛速度。
英文摘要:
Deep reinforcement learning (DRL) has achieved strong performance across a wide range of continuous-control problems. These continuous-control policies, however, are often defined as direct mappings from an observed state to an action or action distribution, requiring a single feed-forward network to construct an optimal control decision in one pass. While effective, this formulation leaves little opportunity for the policy to reconsider or progressively improve an action once an initial prediction has been formed. In this work, we explore an alternative approach: rather than learning only to directly predict an action, can a policy learn to iteratively improve one, and can this iterative process provide advantages during policy learning? We introduce Iterative Action Refinement (IAR), an iterative action-construction method that constructs control actions through a sequence of learned residual corrections. Starting from an initial proposal, a shared refinement network repeatedly conditions on the observed state and the current action proposal, allowing each refinement step to revise the action constructed by preceding steps. The final refined proposal is then used to determine the action executed by the agent. We integrate this iterative action-construction mechanism with Proximal Policy Optimization (PPO), yielding REFINEPPO. We evaluate REFINEPPO across 14 benchmark control tasks, complemented by controlled ablations of refinement depth and update schedules and analyses aimed at understanding why iterative refinement is effective. Across these environments, REFINEPPO matches or exceeds the performance of standard PPO while demonstrating faster convergence on several tasks.