发表机构
Imperial College London(伦敦帝国学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨阶段结构强化学习中单一共享策略与多策略的优劣,提出基于机制持续时间的阶段分解方法,并通过实验验证四个关键假设。
AI 中文摘要
许多强化学习(RL)问题是非平稳但结构化的,可以分解为多个阶段,每个阶段具有各自的转移概率和奖励函数。当阶段序列已知时,常见解决方案是通过增加状态信息以满足马尔可夫性质,并应用标准强化学习技术。然而,先前研究发现,针对不同阶段的多策略方法可以优于在各阶段间共享的单一状态增强策略,其原因尚不明确。在本工作中,我们首先证明共享策略在理论上可以达到任何多策略解决方案的性能。然而,多策略解决方案在实践中是否优于相应的单一共享策略,取决于函数逼近、学习和优化过程,以及对于多策略解决方案而言,样本效率和策略间连续性的损失。我们提出一种基于机制(regime)的阶段分解方法,以识别哪种策略能提供更好的性能。该方法基于瞬态系统动态持续时间相对于准平稳时期持续时间的考量。我们使用不同的非平稳强化学习问题进行数值实验,以验证我们的四个主要假设:(a)较长的阶段持续时间有利于多策略,(b)阶段间的异质性增加了单一策略的负担,(c)多策略需要每个阶段有足够的数据,以及(d)阶段间环境特定的转移动态会影响哪种策略更优。
英文摘要
Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution augments the state with information to satisfy the Markovian property and applies standard RL techniques. However, prior work finds that the multi-policy approach for different phases can outperform a single state-augmented policy shared among the phases, for reasons that remain unclear. In this work, we first show that the shared policy can theoretically achieve performance of any multi-policy solution. However, whether a multi-policy solution can perform better than the corresponding single shared policy in practice depends on function approximation, learning and optimization processes, as well as, for multi-policy solutions, the sample efficiency and loss of continuity from one policy to another. We propose a regime-based phase decomposition method to identify which policy can provide better performance. The method is based on consideration of the duration of transient system dynamics relative to the duration of the quasi-stationary period. Numerical experiments are conducted with different non-stationary RL problems to validate our four major hypotheses: (a) longer phase durations favor multi-policies, (b) the heterogeneity between phases increases the burden on single policy, (c) multi-policies need sufficient data for each phase, and (d) environment-specific transition dynamics between phases can affect which policy is preferable.
Comments40 pages, 13 figures, main paper with appendix