局部引导的演员-评论家:利用感知子目标的评论家训练目标条件演员
Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware Critic
浏览论文内容
中文总结 AI 辅助
针对目标条件强化学习长 horizon 稀疏奖励的问题,该文提出 LG-AC 方法,以子目标条件值函数总和实现密集事后重标记,在目标链任务中性能优于 RIS、RS 等方法。
中文摘要 AI 辅助
目标条件强化学习在奖励稀疏时处理长 horizon 会遇到困难。虽然规划器可提供子目标以指导低层策略,但在测试时使用它可能会引入实际的子目标管理难题。另一种范式利用高层规划器辅助学习,而策略仅以最终目标为条件,支持无规划器部署。在这些方法中,带想象子目标的强化学习(RIS)引入正则化项,鼓励策略对最终目标采取与中间目标相同的动作。然而,当中间目标是低维时,这种正则化可能导致目标链问题。基于势的奖励 shaping(PBRS)将计划转换为额外奖励,同时确保最优策略保持不变,但它可能在终端状态产生欺骗性奖励。我们研究这些失败案例,首先提出一种替代奖励 shaping 方法(RS),它消除了这些欺骗性奖励,但牺牲了 PBRS 的理论保证。与该 RS 变体类似,我们随后提出另一种名为局部引导的演员-评论家(LG-AC)的方法,该方法因智能体到达中间目标而给予奖励。与 RS 中中间奖励隐含在 shaping 信号中不同,我们明确地将值估计器以完整的中间目标序列为条件,但将值函数表示为子目标条件值函数的总和,从而实现密集的事后重标记。我们在具有挑战性的目标链要求的任务中评估所有这些方法,并从经验上强调,在某些情况下,动作正则化或奖励 shaping 会产生低性能,而 LG-AC 在所有任务中实现最佳整体性能。
英文摘要
Goal-conditioned reinforcement learning struggles with long horizons when rewards are sparse. While a planner can provide subgoals to guide a low-level policy, its use at test time may introduce practical subgoal management difficulties. An alternative paradigm utilizes a high-level planner to assist learning, while the policy remains conditioned only on the final goal, enabling planner-free deployment. Among these methods, Reinforcement Learning with Imagined Subgoals (RIS) introduces a regularization term that encourages the policy to take the same actions for the final goal as it does for an intermediate goal. This regularization, however, may lead to goal-chaining issues when intermediate goals are low-dimensional. Potential-based reward shaping (PBRS) translates plans into an additional reward while ensuring that the optimal policy remains unchanged. Yet, it can generate deceptive rewards in terminal states. We study these failure cases and first propose an alternative reward shaping method (RS) that removes these deceptive rewards at the expense of theoretical guarantees of PBRS. Similar to this RS variant, we then propose another method named Locally-Guided Actor Critic (LG-AC) that rewards the agent for reaching intermediate goals. Unlike RS, where intermediate rewards are implicit in the shaping signal, we explicitly condition a value estimator on the full sequence of intermediate goals but represent the value function as a sum of subgoal-conditioned value functions, enabling dense hindsight relabeling. We evaluate all these methods in tasks with challenging goal-chaining requirements and empirically highlight specific cases in which either action regularization or reward shaping yield low performance, while LG-AC achieves the best overall performance across tasks.
发表机构
- ISIR
- Sorbonne Université(索邦大学)
- CNRS(法国国家科学研究中心)
机构由 AI 辅助整理,请以论文原文为准。