方差感知的在线强化学习细粒度间隙依赖界
Variance-Aware Fine-Grained Gap-Dependent Bounds for Online Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本文为免模型在线强化学习提出方差感知的细粒度间隙依赖界,通过改进的UCB-Bernstein+算法同时优化遗憾与局部切换成本,并给出首个此类上界及实证验证。
中文摘要 AI 辅助
我们研究分幕表格马尔可夫决策过程的免模型在线强化学习(RL),重点关注间隙依赖的遗憾和策略切换成本。虽然针对使用Hoeffding型探索奖励的免模型RL算法已建立了细粒度间隙依赖分析,但对于使用基于方差的探索奖励的免模型算法,此类结果仍然未知,尽管它们在最坏情况和粗粒度间隙依赖保证方面表现更优。在本文中,我们通过为UCB-Bernstein+(一种改进的UCB-Bernstein算法)建立首个方差感知免模型在线RL中的细粒度间隙依赖遗憾上界,解决了这一开放问题。此外,通过将阶段式策略更新设计整合到我们的细粒度框架中,并使用改进的基于方差的奖励,我们实现了迄今为止已知最佳的间隙依赖局部切换成本。另外,我们的分析在遗憾和局部切换成本方面,相较于原始UCB-Bernstein算法,改进了最坏情况保证。数值实验进一步表明,UCB-Bernstein+在遗憾和局部切换成本方面均取得了良好的实证性能。
英文摘要
We study model-free online reinforcement learning (RL) for episodic tabular Markov decision processes, focusing on both gap-dependent regret and policy switching cost. While fine-grained gap-dependent analysis has been established for model-free RL algorithms using Hoeffding-type exploration bonuses, such results for model-free algorithms with variance-based exploration bonuses remain unknown, despite their superior worst-case and coarse-grained gap-dependent guarantees. In this paper, we resolve this open problem by establishing the first fine-grained gap-dependent regret upper bound for UCB-Bernstein+, a refined UCB-Bernstein algorithm, in variance-aware model-free online RL. Moreover, by integrating a stage-wise policy update design into our fine-grained framework and using refined variance-based bonuses, we achieve the best-known gap-dependent local switching cost to date. In addition, our analysis yields improved worst-case guarantees for both regret and local switching cost over the original UCB-Bernstein algorithm. Numerical experiments further demonstrate that UCB-Bernstein+ achieves favorable empirical performance in both regret and local switching cost.
发表机构
- The Pennsylvania State University(宾夕法尼亚州立大学)
- University of Notre Dame(圣母大学)
机构由 AI 辅助整理,请以论文原文为准。