arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从连通性到奖励:基于有向状态图的稠密奖励学习

From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs

Shuyuan Zhang, Zihan Wang, Xiao-Wen Chang, Doina Precup

arXiv 2609.10781首次发表:更新:

发表机构

McGill University; Mila – Quebec AI Institute; Google DeepMind(麦吉尔大学; 米拉–魁北克人工智能研究所; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出G2QDR框架,利用有向状态图预测状态连通性并转化为稠密奖励,以增强目标条件分层强化学习在拟度量环境中的性能。

AI 中文摘要

图与目标条件分层强化学习(GCHRL)的集成日益受到关注,因为图天然地编码了任务层次结构,有助于有效的子目标采样。然而,现有方法往往忽视内在的连通性信息,未能充分利用底层拓扑结构来实现高效学习。大多数基于图的GCHRL方法将图用作随机采样工具,而非编码连通性和状态可达性信息的环境模型。这一局限在拟度量环境中尤为突出,在这种环境中,状态转换固有的不对称性对稳定的策略学习和鲁棒的路径规划构成了根本性挑战。在本文中,我们通过引入一种状态连通性模型来解决这些问题,该模型旨在预测不对称环境中成对状态之间的连通强度。我们将这些连通强度转化为标量辅助稠密奖励,在多个层次级别上提供连续指导。我们证明,我们提出的框架——图引导拟度量稠密奖励(G2QDR)——理论上可以集成到任何现有的GCHRL架构中,并且状态连通性模型通过一个在探索过程中生成的有向状态图上训练的神经网络高效实现。在广泛的稀疏奖励环境中的实证结果表明,总体而言,G2QDR能够在可接受的计算开销下提升基线GCHRL方法的性能。

英文摘要

The integration of graphs with Goal-Conditioned Hierarchical Reinforcement Learning (GCHRL) has received increasing attention, as graphs naturally encode task hierarchies for effective subgoal sampling. However, existing methods often overlook intrinsic connectivity information, failing to fully leverage the underlying topology for efficient learning. Most graph-based GCHRL methods use the graph as a stochastic sampling tool rather than as an environmental model that encodes connectivity and state-accessibility information. This limitation is particularly acute in quasimetric environments, where the inherent asymmetry of state transitions poses a fundamental challenge to stable policy learning and robust path planning. In this paper, we address these problems by introducing a state connectivity model designed to predict pairwise state connectivity strength in asymmetric environments. We transform these connectivity strengths into scalar auxiliary dense rewards, providing continuous guidance across multiple hierarchical levels. We demonstrate that our proposed framework, Graph-Guided Quasimetric Dense Reward (G2QDR), can theoretically be integrated into any existing GCHRL architecture, and the state connectivity model is efficiently implemented via a neural network trained on a directed state graph generated during exploration. Empirical results across a wide range of sparse reward environments indicate that, in general, G2QDR can enhance the performance of baseline GCHRL approaches with acceptable computational overhead.

Journal refZhang, S., Wang, Z., Chang, X.-W., & Precup, D. (2026). From connectivity to rewards: Dense reward learning with directed state graphs. Transactions on Machine Learning Research. https://openreview.net/forum?id=F65zrefsjB

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑