arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12135cs.AI

用于分层强化学习的Q形选项

Q-Shaped Options for Hierarchical Reinforcement Learning

发表机构牛津大学
查看机构详情
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Clarisse Wibault, Antoine Gorceix, Antonio Léon Villares, Alexey Zakharov, Evangelos Chatzaroulas, Michael Matthews, Eduardo Pignatelli, Jakob Foerster

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对分层强化学习中动作抽象的缺陷,提出Q形选项算法,满足动作抽象的三个必要条件,在离线目标条件的运动和操作环境中优于基线算法,在其他算法失效的任务中实现非零性能。

中文摘要 AI 辅助

学习解决长 horizon、目标条件任务需要智能体在扩展时间尺度上推理,并在广泛状态范围内行动。原则上,分层强化学习(Hierarchical Reinforcement Learning, HRL)通过动作(时间)与状态(空间)抽象的交互应对这两个挑战:其一,使用动作抽象将时间扩展行为表示为选项,可降低有效决策 horizon;其二,在决策过程的每个层级启用不同状态抽象,能为学习实现更优的数据聚合。然而,分层策略这两大优势的实现依赖于学习合适的动作抽象,当前HRL算法存在两类缺陷:部分算法丢弃了最优控制所需的选项间区分,彻底破坏了层级结构;另一部分算法保留了不必要的区分,虽维持了horizon缩减,却放弃了更粗粒度的状态抽象。本研究明确了动作抽象的三个 desiderata( desiderata 为“ desideratum”的复数,指“必要条件”),并引入Q形选项(Q-Shaped Options, QSO)以满足全部三个条件。QSO基于一种架构构建,该架构在分层的每个层级拥有独立的状态值函数、Q函数和策略;它将连续层级间的动作抽象学习为由对应Q函数塑造的共享编码器,其中低层Q函数将选项作为目标,促使抽象保留最优控制所需的区分,高层Q函数则将选项作为动作,促使丢弃不必要的区分。在离线目标条件的运动和操作环境中,QSO学习到语义有意义的选项空间,且优于基线算法,在所有其他评估算法均失败的任务中实现了非零性能。

英文摘要

Learning to tackle long-horizon, goal-conditioned tasks requires an agent to reason over extended timescales and act across a broad range of states. In principle, Hierarchical Reinforcement Learning (HRL) addresses both challenges through the interaction between action (temporal) and state (spatial) abstraction. First, using an action abstraction to represent temporally extended behaviour as options reduces the effective decision horizon. Second, enabling different state abstractions at each level of the decision process permits greater data aggregation for learning. However, realising these two benefits of a hierarchical policy depends on learning an appropriate action abstraction. Current HRL algorithms fail in one of two ways. Some discard distinctions between options needed for optimal control, undermining hierarchy altogether. Others retain unnecessary distinctions, preserving horizon reduction, but forfeiting coarser state abstraction. In this work, we characterise three desiderata for an action abstraction. We introduce Q-Shaped Options (QSO) to address all three. QSO builds on an architecture with distinct state-value functions, Q functions and policies at each level of the hierarchy. It learns the action abstraction between consecutive levels as a shared encoder shaped by their respective Q functions. The low-level Q function uses the option as a goal, encouraging the abstraction to retain distinctions necessary for optimal control. The high-level Q function uses it as an action, encouraging unnecessary distinctions to be discarded. Across offline goal-conditioned locomotion and manipulation environments, QSO learns semantically meaningful option spaces and outperforms baselines, achieving non-zero performance in tasks where all other evaluated algorithms fail.

↑