发表机构
The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对离线目标条件强化学习中稀疏奖励和长时程依赖导致的奖励传播延迟问题,提出RSIQL方法,通过在进展中间状态引入额外奖励信号,以简单非分层结构提升性能,实验验证其优于目标条件IQL并媲美分层方法。
AI 中文摘要
目标条件强化学习旨在学习能够达到指定目标的策略,但在稀疏奖励和长时程依赖的离线设置中仍然具有挑战性。在此类设置中,目标完成信息在时间上可能远离促成成功的早期决策,而离线价值估计又会引入额外误差。我们从奖励传播的角度研究这一问题,并在一个简化的延迟目标设置中表明,目标导向的价值分离相对于局部估计误差可能变得很小。受此分析启发,我们提出了奖励刺激隐式Q学习(RSIQL),一种简单的非分层方法,它在离线轨迹中取得进展的中间状态处引入额外的奖励信号。RSIQL使用一个辅助的目标条件价值函数来识别估计能向目标取得进展的中间状态,并应用奖励刺激以提供延迟更少的训练监督。与分层方法不同,RSIQL不学习单独的高层子目标策略。在D4RL目标到达基准和OGBench上的实验表明,RSIQL平均优于目标条件IQL,并达到与分层离线目标条件方法竞争的性能,同时保持简单的平面策略结构。
英文摘要
Goal-conditioned reinforcement learning aims to learn policies that reach specified goals, but remains challenging in offline settings with sparse rewards and long-horizon dependencies. In such settings, goal-completion information can be temporally distant from the early decisions that enable success, while offline value estimation introduces additional error. We study this issue from a reward-propagation perspective and show, in a stylized delayed-goal setting, how goal-directed value separation can become small relative to local estimation error. Motivated by this analysis, we propose Reward Stimulation Implicit Q-Learning (RSIQL), a simple non-hierarchical method that introduces additional reward signals at progress-making intermediate states in offline trajectories. RSIQL uses an auxiliary goal-conditioned value function to identify intermediate states estimated to make progress toward the goal and applies reward stimulation to provide less-delayed training supervision. Unlike hierarchical methods, RSIQL does not learn a separate high-level subgoal policy. Experiments on D4RL goal-reaching benchmarks and OGBench show that RSIQL improves over goal-conditioned IQL on average and achieves performance competitive with hierarchical offline goal-conditioned methods, while retaining a simple flat policy structure.