Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning
在线强化学习中的非渐近最优策略识别保证
机构 * Imperial College London(伦敦帝国学院) ; Boston University(波士顿大学) ; Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所)
AI总结 研究在线表格强化学习中的最优策略识别问题,通过为导航与停止(NaS)算法提供非渐近样本复杂度保证,揭示其样本复杂度不仅依赖特征时间,还与MDP连通性等有关,填补了相关空白。
Comments 64 pages, 2 figures