AI 中文总结
本文识别分层强化学习中执行与策略次优性,提出统一价值函数和广义贝尔曼方程,实现独立执行选择与改进,实验验证互补收益。
AI 中文摘要
分层强化学习利用时间上延展的子任务进行探索,然而承诺执行这些子任务可能会限制部署和策略学习。我们识别并区分了由此产生的执行次优性和策略次优性。任务树和执行树将奖励目标与策略选择及决策中断区分开来。针对HRL的统一价值函数和四阶段广义分层贝尔曼方程支持对这两种损失的共同分析。在有界奖励和均匀终止条件下,我们建立了分层策略和执行改进结果。在其余节点策略固定的情况下,任务子树兼容性和原始执行模式下的节点策略最优性确定了马尔可夫执行何时是最优的。由此产生的分解导致了行为、目标和部署的独立执行选择。我们通过执行改进以及任意层级深度下的一阶段或两阶段策略改进来实例化这一原则。基于选项和目标条件的实验证明了改变执行和改变学习目标所带来的互补收益。受控随机环境展示了这些收益如何依赖于随机转移强度和空间结构。该框架使执行设计成为分层策略优化的显式组成部分。
英文摘要
Hierarchical reinforcement learning uses temporally extended subtasks for exploration, yet committing to their execution can restrict both deployment and policy learning. We identify and separate the resulting execution and policy suboptimality. Task and execution trees distinguish reward objectives from policy choices and decision interruption. A Unified Value Function for HRL and a four-stage Generalized Hierarchical Bellman Equation then support a common analysis of both losses. Under bounded rewards and uniform termination, we establish hierarchical policy and execution improvement results. With the remaining node policies fixed, task-subtree compatibility and node-policy optimality under the original execution mode establish when Markov execution is optimal. The resulting decomposition leads to independent execution choices for behavior, targets, and deployment. We instantiate this principle through execution improvement and one-stage or two-stage policy improvement at arbitrary hierarchy depth. Option-based and goal-conditioned experiments demonstrate complementary gains from changing execution and changing the learning target. Controlled stochastic environments show how these gains depend on stochastic transition strength and spatial structure. This framework makes execution design an explicit component of hierarchical policy optimization.