GTRL:基于时间差分的基础分治价值学习
GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences
浏览论文内容
中文总结 AI 辅助
GTRL通过将一步TD目标融入分治更新,解决了离线GCRL中随机动态下的估值偏差和状态-目标对无更新问题,在OGBench任务上取得最高平均成功率。
中文摘要 AI 辅助
在离线目标条件强化学习(GCRL)中,分治方法通过在一个子目标处连接两个较短的片段来扩展到长视野。然而,在随机动态下,该规则的基础情况会通过数据对最幸运的轨迹进行估值。子目标还必须位于一条共享轨迹上,因此,如果没有轨迹连接某个状态-目标对,该对将完全得不到价值更新。为了解决这两个问题,我们提出了有基础传递强化学习(GTRL),一种离线GCRL价值学习算法,它通过一步TD目标来夯实分治更新。在单步中,TD是正确的,因为其目标对后继状态进行平均,且不需要子目标。GTRL将该目标添加到组合中而不是替换它,因此每一对都能收到更新,且组合仍然承载长视野。GTRL还通过根据每个目标从其他后继状态的可达性对其重新加权,来纠正事后重标记带来的偏差。我们在跨越随机、确定性和拼接环境的十九个OGBench任务上评估了我们的算法,它取得了最高的平均成功率。代码即将发布。
英文摘要
In offline goal-conditioned reinforcement learning (GCRL), divide-and-conquer scales to long horizons by joining two shorter segments at a subgoal. However, under stochastic dynamics, the base case of this rule values the luckiest trajectories through the data. The subgoal must also lie on a shared trajectory, so a state-goal pair that no trajectory connects gets no value update at all. To address both, we present Grounded Transitive RL (GTRL), an offline GCRL value learning algorithm that grounds the divide-and-conquer update with a one-step TD target. Over a single step, TD is correct, as its target averages over the successors and needs no subgoal. GTRL adds this target to the composition rather than replacing it, so every pair receives an update, and the composition still carries the long horizon. GTRL also corrects the bias from hindsight relabeling by reweighting each goal against how reachable it was from other successors. We evaluate our algorithm on nineteen OGBench tasks spanning stochastic, deterministic, and stitching environments, where it achieves the highest average success rate. Code will be released soon.