AI 中文总结
针对离线策略分层强化学习中子目标重标记导致的价值更新偏差问题,提出HTB方法,通过时间感知引导、多步回报及非负残差约束,在AntFall任务上显著优于HIRO。
AI 中文摘要
离线策略分层强化学习必须在低层策略变化的同时估计高层决策的价值。HIRO通过子目标重标记来调整回放数据,但标签改变后,价值更新针对的是重标记后的子目标,而非高层策略原本需要更新的子目标。我们提出分层时间感知引导(HTB),该方法在当前低层策略下评估指定子目标,同时保留累积的任务奖励。剩余执行时间区分了子目标的延续与新的高层决策。结合原始动作条件,它能够基于平稳的环境转移规律进行离线策略贝尔曼更新。HTB将这些单步更新与多步后缀回报及截断重标记相结合,减少了对中间价值估计的依赖。共享价值组件支持跨动作学习,而非负残差则约束了相对于该组件的向上修正。在固定的混合权重0.95下,在10M环境步数内,HTB在五个配对种子上达到32.8%的AntFall成功率,而匹配的局部HIRO仅为9.6%。消融实验识别了递归延续和混合监督的贡献;固定策略测试显示,对于其回报被排除在拟合之外的动作,预测更为准确。
英文摘要
Off-policy hierarchical reinforcement learning must estimate the values of high-level decisions while the low-level policy changes. HIRO adapts replay data through subgoal relabeling, but after a label change, the value update targets the relabeled subgoal instead of the subgoal the high-level policy originally needed to update. We propose Hierarchical Time-aware Bootstrapping (HTB), which evaluates specified subgoals under the current low-level policy while retaining accumulated task rewards. Remaining execution time distinguishes subgoal continuation from a new high-level decision. Together with primitive-action conditioning, it enables off-policy Bellman updates based on the stationary environment transition law. HTB combines these one-step updates with multi-step suffix returns and truncated relabeling, reducing dependence on intermediate value estimates. A shared value component supports learning across actions, while nonnegative residuals constrain upward corrections relative to that component. At a fixed mixture weight of 0.95, HTB achieves 32.8% AntFall success versus 9.6% for matched local HIRO over five paired seeds at 10M environment steps. Ablations identify contributions from recursive continuation and mixed supervision; fixed-policy tests show more accurate predictions for actions whose returns were excluded from fitting.