发表机构
Imperial College London; Boston University; Broad Institute of MIT and Harvard(伦敦帝国学院; 波士顿大学; MIT和哈佛大学Broad研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在线表格强化学习中的最优策略识别问题,通过为导航与停止(NaS)算法提供非渐近样本复杂度保证,揭示其样本复杂度不仅依赖特征时间,还与MDP连通性等有关,填补了相关空白。
AI 中文摘要
在这项工作中,我们研究在线表格强化学习中的最优策略识别(BPI)问题。这是一个主动序贯假设检验问题,学习者的目标是以高置信度识别马尔可夫决策过程(MDP)中的最优策略,同时最小化预期样本复杂度。我们考虑具有确定性奖励的在线设置,智能体必须策略性地在MDP中导航以有效探索。先前文献为BPI提供了渐近最优方法,如导航与停止(NaS)算法及其变体,但现有分析仍是渐近的。我们通过为NaS提供首个非渐近样本复杂度保证来填补这一空白,表明其样本复杂度不仅取决于特征时间,还取决于基础MDP的连通性、最优特征时间的曲率以及其他依赖实例的量。我们识别出这些额外属性并明确它们对整体样本复杂度的贡献。
英文摘要
In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning. This is an active sequential hypothesis testing problem in which the learner's objective is to identify an optimal policy in a Markov Decision Process (MDP) with high confidence, while minimizing the expected sample complexity to do so. We consider an online setting with deterministic rewards, where the agent must strategically navigate through the MDP in order to effectively explore. Previous works in the literature have provided asymptotically optimal methods for BPI, such as the Navigate and Stop (NaS) algorithm and its variants, however existing analysis remains asymptotic. In this work, we fill that gap by providing the first non-asymptotic sample complexity guarantees for NaS, showing that its sample complexity depends not only on the characteristic time, but also on the connectivity of the underlying MDP, the curvature of the optimal characteristic time, and other instance-dependent quantities. We identify these additional attributes and make explicit their contributions to the overall sample complexity.
Comments64 pages, 2 figures