学习从有限时间可达性中获取目标到达的拟度量几何
Learning Goal-Reaching Quasimetric Geometry From Finite-Time Reachability
- University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对目标条件强化学习,提出ReQRL方法,利用有限时域可达性约束价值梯度,解耦动力学可达性与边界几何,在OGBench上表现优于或媲美现有方法。
AI中文摘要:
在目标条件强化学习(GCRL)中,拟度量学习将目标到达成本建模为拟度量距离,将局部约束与全局价值几何联系起来。然而,其局部约束应反映有限时域内控制组合的方向依赖效应以及环境可行性。我们提出ReQRL,通过有限时域可达性约束评论家的价值梯度。借鉴状态约束最优控制,我们将动力学可达性与边界几何解耦,并从数据中估计两者。在OGBench上,我们的方法优于或可与现有拟度量方法及其他离线GCRL方法相媲美。
英文摘要:
In goal-conditioned reinforcement learning (GCRL), quasimetric learning models goal-reaching costs as quasimetric distances, connecting local constraints to global value geometry. Its local constraints, however, should reflect the direction- dependent effects of control composition over a finite horizon together with environmental feasibility. We propose ReQRL, which constrains the critic's value gradients through finite-horizon reachability. Drawing on state-constrained optimal control, we decouple dynamical reachability from boundary geometry, estimating both from data. On OGBench, our method outperforms or rivals existing quasimetric approaches and other offline GCRL methods.