AI 中文总结
针对现有双层强化学习算法可扩展性差或样本复杂度高的问题,提出基于玻尔兹曼策略最优性的无Hessian超梯度双层RL算法,实现更优的迭代与样本复杂度,且无需外层目标的PL条件假设。
AI 中文摘要
双层强化学习(RL)是强化学习领域中的重要框架,可用于形式化各类问题,如元学习、分层任务分解以及人类反馈强化学习(RL-HF)。大多数双层强化学习算法要么因使用含Hessian的超梯度而不具备可扩展性,要么因采用基于惩罚的近似方法而存在样本复杂度高的问题。在本研究中,我们针对熵正则化折扣RL目标函数,利用玻尔兹曼策略的最优性,提出了一种基于超梯度的双层RL算法。我们提出的算法无需Hessian,在温和的正则性条件下,可达到$O(ε^{-1})$的迭代复杂度以及$\tilde{O}(ε^{-2})$的当前最优样本复杂度。此外,在我们的收敛性分析中,我们能够移除现有最优样本复杂度研究中对外层目标函数施加的Polyak-Lojasiewicz(PL)条件假设。
英文摘要
Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories of problems, such as meta-learning, hierarchical task decomposition, and reinforcement learning from human feedback (RL-HF). Most of the bilevel RL algorithms are either not scalable because of using hypergradient with Hessian, or they suffer from high sample complexity because of using penalty-based approximation methods. In this work, we propose a hypergradient-based bilevel RL algorithm using the optimality of the Boltzmann policy for the entropy regularized discounted RL objective function. Our proposed algorithm is Hessian-free and obtains an iteration complexity of $O(ε^{-1})$ and state-of-the-art sample complexity of $\tilde{O}(ε^{-2})$ under mild regularity conditions. Further, in our convergence analysis, we are able to remove the assumption of the Polyak-Lojasiewicz (PL) condition on the outer-level objective function present in the prior state-of-the-art sample complexity work.