发表机构
Universidade Federal de Minas Gerais; University of Alberta; Alberta Machine Intelligence Institute (Amii); Canada CIFAR AI Chair(米纳斯吉拉斯联邦大学; 阿尔伯塔大学; 阿尔伯塔机器智能研究所; 加拿大CIFAR AI主席)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对离线目标条件强化学习的长时程任务,提出广义隐式时间抽象(GITA),通过多尺度聚合优势加权监督,兼顾局部与长距离信号,在OGBench上显著优于现有基线。
AI 中文摘要
现有的离线目标条件强化学习(GCRL)方法在处理长时程任务时面临困难。折扣会缩小远距离状态之间的价值差异,直至其低于函数逼近误差,使得智能体缺乏对状态进行排序的信号。时间抽象将k个环境步骤视为单个转移,能够恢复长距离上的这种信号,但不存在一个固定的k值适用于所有状态-目标距离:较大的k值能保留长时程时间距离上的价值差异,但会抹平邻近状态之间的区别,而较小的k值则相反。我们将这一权衡明确化,并引入广义隐式时间抽象(GITA),该方法将单个价值函数以k为条件。GITA通过聚合多个k值下的优势加权监督来训练一个策略,因此,为某个状态-目标对分配较大正优势的尺度会对其更新贡献更强。GITA无需在局部分辨率和长距离信号之间做出选择;它无需固定于单一k值即可同时保留两者。在OGBench上,GITA优于广泛的离线GCRL基线,在所有任务上的平均成功率比HIQL提高了25个百分点(相对提升73%)。它也比最强的固定k方法OTA提高了7个百分点(相对提升14%)。
英文摘要
Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).